Selected work

Selected work in AI infrastructure, NLP, and applied research.

In production

Molcajete

AI-powered transcription and analysis pipeline for political focus group research.

A full audio-to-insight pipeline replacing error-prone transcription and hours of note-taking per project. Speaker diarization, transcription, theme classification, and integrated reporting — all surfaced through a tooling layer that researchers actually use.

2,000+
hours of audio processed
<60 min
turnaround per project

Adapta

LLM fine-tuning infrastructure and data preprocessing pipeline for Mexican Spanish political analysis.

Specialized LLMs, produced by a reproducible fine-tuning and evaluation pipeline. Built for empirical comparison of base models and prompts.

40+
evaluation metrics
100+
training runs

Nopalero

Automated participant screening system for qualitative recruitment.

Automated intake pipeline that replaces hours of manual data entry per project. Combines OCR, fraud detection, and socioeconomic classification — so analysts focus on the decisions, not the paperwork.

48
validation checks
0
manual data entry

Azulejo

Turns rough notes into an editable, on-brand PowerPoint at the press of a button.

The brand template is codified once, so every slide comes out compliant by construction. Content lives in plain text and the slides regenerate from it, versionable and reproducible like code. The output is a standard .pptx that anyone can edit, with no platform in the loop.

94
layout templates
<5 s
to rebuild 50 slides

Open source

Python scraper and parser for Brazilian Supreme Court (STF) case data.

Typer-based CLI with three cache-first stages — scrape, download, extract — supporting heavily sharded runs with proxy rotation, feeding a DuckDB warehouse. Multiple OCR backends, including self-hosted Tesseract on fly.io for cheap inference.

R$ 52
yearly HC sweep
0.93/s
PDFs sustained
0.28/s
cases sustained
4
OCR backends

CLI for an AI agent to keep a Zotero library in order.

Checks tags, metadata, and collections; enriches metadata from Crossref and arXiv; compresses PDFs. Implements a validation loop with backups and write-ahead logs — making it safe to let an AI agent touch the database.

Have a problem that doesn't fit a template?

Most of the work above started as someone saying exactly that.

Start a conversation