AI, Data & UX-Driven Product Growth
I build AI features — LLM flows, recommendation models, AI-powered insights — and wrap them in evals so you know they actually work, and keep getting better. Not vibes. Measured. (I ship my own AI products too — ViziAI, Salta, Etko.)
From a first prototype to a production feature your users rely on.
Chat, generation, extraction, summarization, and agentic flows — built into your product.
Matching, ranking, personalization, and scoring models trained on your signals.
Turn raw, messy data into clear, actionable output people trust.
Tools that do the task — not just describe it — with humans in the loop where it counts.
Most AI ships on vibes: someone eyeballs a few outputs and calls it done. Evals replace that with a measurable, repeatable test — does the output meet the bar, at scale? It's the difference between "seems fine" and "98% compliant, and I can prove the trend."
The model produces an output — a summary, a match, a notification.
A script (or an LLM judge) scores it against the bar — accuracy, format, safety.
Failures get a second pass with a stronger model — cheap-first, escalate only on miss.
Still failing? Route to a person. Nothing broken reaches the user.
The metric you care about drives everything — and every release is re-evaluated, not shipped and forgotten.
A clear path — you always know what's happening and why.
We pin down the user problem and the one metric success moves.
A fast, working version — often rules-first — to prove the value.
I build the tests that measure quality, so we ship on evidence.
Live with fallbacks and monitoring — then improve on the numbers.
A few AI/ML systems I built and measured in production.
Raw error codes → LLM-generated, actionable descriptions. Killed 30–40 min of per-error research; beta customers sent unsolicited thank-yous.
ML recommendation engine surfacing best-fit candidates, validated with an MLOps eval loop (bi-weekly retraining, tracked accuracy).
LLM notification compliance, via an eval → reflection → human-fallback flow — at −75% cost by escalating only on failure.
Diagnosed a reliability collapse (usage 90%→10%), designed tiered logic (85% accuracy on 10 dimensions), tripled adoption back to 30%.
Tell me what you're trying to build or measure. I'll reply with whether I can help and what the first step looks like.