Onur Ovalı

AI, Data & UX-Driven Product Growth

AI Product & Evals

Ship AI you can trust.

I build AI features — LLM flows, recommendation models, AI-powered insights — and wrap them in evals so you know they actually work, and keep getting better. Not vibes. Measured. (I ship my own AI products too — ViziAI, Salta, Etko.)

Work with me → How it works
What I build

AI that does real work.

From a first prototype to a production feature your users rely on.

01

LLM features

Chat, generation, extraction, summarization, and agentic flows — built into your product.

02

Recommendation & ML

Matching, ranking, personalization, and scoring models trained on your signals.

03

AI-powered insights

Turn raw, messy data into clear, actionable output people trust.

04

Agents & automation

Tools that do the task — not just describe it — with humans in the loop where it counts.

The difference

What evals are — and why they matter.

Most AI ships on vibes: someone eyeballs a few outputs and calls it done. Evals replace that with a measurable, repeatable test — does the output meet the bar, at scale? It's the difference between "seems fine" and "98% compliant, and I can prove the trend."

Read my deep-dive on evals →

How an eval gate works
1

Generate

The model produces an output — a summary, a match, a notification.

2

Auto-eval

A script (or an LLM judge) scores it against the bar — accuracy, format, safety.

3

Reflect

Failures get a second pass with a stronger model — cheap-first, escalate only on miss.

4

Human fallback

Still failing? Route to a person. Nothing broken reaches the user.

How I implement

Business metric first, then the model.

The metric you care about drives everything — and every release is re-evaluated, not shipped and forgotten.

Business metricwhat to improve Buildrules first, then AI Eval harnessscore every output Shipwith fallbacks Monitortrack the trend ↻ CONTINUOUS — EVERY RELEASE RE-EVALUATED
How we’d work together

From idea to a feature you trust.

A clear path — you always know what's happening and why.

1

Scope

We pin down the user problem and the one metric success moves.

2

Prototype

A fast, working version — often rules-first — to prove the value.

3

Eval harness

I build the tests that measure quality, so we ship on evidence.

4

Ship & iterate

Live with fallbacks and monitoring — then improve on the numbers.

Proof

Real results, shipped.

A few AI/ML systems I built and measured in production.

NPAW · analytics SaaS
30 min → seconds

Raw error codes → LLM-generated, actionable descriptions. Killed 30–40 min of per-error research; beta customers sent unsolicited thank-yous.

Kariyer.net · marketplace
15 days → 3 days

ML recommendation engine surfacing best-fit candidates, validated with an MLOps eval loop (bi-weekly retraining, tracked accuracy).

Etko · consumer app
50% → 98%

LLM notification compliance, via an eval → reflection → human-fallback flow — at −75% cost by escalating only on failure.

NPAW · reliability
AI usage

Diagnosed a reliability collapse (usage 90%→10%), designed tiered logic (85% accuracy on 10 dimensions), tripled adoption back to 30%.

Let's talk

Work with me.

Tell me what you're trying to build or measure. I'll reply with whether I can help and what the first step looks like.

or email me directly at hi@onurovali.me