New Roles Series
Evaluation, Search / RAG, and Agents — the roles every team is hiring for and few people can teach. Real numbers, a real case study in every chapter, and labs you run free. Not blog-post hand-waving.
Labs run free, no API key · One-time purchase · Learn online
Why it's different
Position bias isn't just defined — it's measured: 8/40 pairs flipped = 20%. You see the failure, then the fix.
Each chapter opens with a real, named scenario — the mistake, the metric, the fix — and one company you follow across the whole book, from first demo to shipped system.
Start with seven free Python labs covering evaluation, retrieval, and agent stopping rules. All run offline, with no API key; the judge lab also offers an optional paid model mode. See the free labs on GitHub →
Evaluation labs · Search / RAG lab · Agent lab. Small, deterministic examples with regression checks; retrieval and agents use simulations.
Scenario diagnosis, not recall — the way a real interview or design review would push you. Every answer explains the why.
The courses
Each role comes in three editions — for working engineers, for students & new grads, and for managers & founders who buy but don't build. Same hands-on method, pitched to you.
Measure systems that don't give the same answer twice — the discipline behind shipping LLMs you can trust.
Build retrieval that makes language models tell the truth about your own documents — and prove it with numbers.
Build LLM agents that finish the task — reliably, cheaply, and without taking an action they can't take back.
Read any Chapter 1 free right now — no signup. All three AI Evaluation Engineer editions are available now; the RAG and Agent courses are on their way.
The free chapters are open above — no email needed. If you'd like the complete course as a PDF and a heads-up the day checkout opens, drop your email. That's the only time we'll write.
New-role drops & launch discounts. No spam, unsubscribe anytime.