AI Engineer Dojo All editions
Managers, founders & PMs · The AI Evaluation Engineer

Evaluating Your AI

You don’t write the evals — but you must know whether your team’s are any good, when to trust a number, and who to hire.

  1. Ch 1

    Why "It Works in the Demo" Will Burn You

    The single most expensive mistake in AI is trusting a demo. A demo is a handful of cherry-picked inputs on a system whose outputs change every time…

    Free preview
  2. Ch 2

    What a Healthy Eval Setup Looks Like

    You don't need to build the system — but you should recognize one that's working. A healthy setup has two halves: a gate before you ship, and a…

    Free preview
  3. Ch 3

    The Dataset Is the Asset

    Teams obsess over which model and which prompt. The thing that actually determines whether your evaluation tells the truth is the dataset — and it's…

    Free preview
  4. Ch 4

    Metrics Without the Math

    You'll hear precision, recall, F1, faithfulness. You don't need the formulas — you need to know which one maps to your risk, because the choice is a…

    Free preview
  5. Ch 5

    When a Model Grades Your Model

    For subjective quality — "is this reply helpful and on-brand?" — there's no answer key, so teams use a strong model as the grader ("LLM-as-judge").…

    Free preview
  6. Ch 6

    Buying Human Judgment

    Automated scores ultimately rest on human judgment. Knowing when to spend on humans — and how to tell if that spend is producing reliable labels —…

    Free preview
  7. Ch 7

    Evaluating RAG / "Chat With Your Docs"

    If you're building "chat with your documents," "ask our knowledge base," or any assistant grounded in your data, you're building RAG — and it fails…

    Free preview
  8. Ch 8

    Why Agents Are a Different Risk

    "Agents" — AI that takes multiple steps and uses tools (sends email, queries databases, takes actions) — are the most exciting and most dangerous…

    Free preview
  9. Ch 9

    Safety: Managing the Worst Case

    Quality is about the average experience. Safety is about the worst one — the output that makes the news, triggers a lawsuit, or leaks a customer's…

    Free preview
  10. Ch 10

    Catching Silent Regressions in Production

    The scariest AI failure makes no noise. Nothing crashes, no error logs, no alert — quality just quietly drops, and the first signal is angry…

    Free preview
  11. Ch 11

    Building an Eval Culture

    The best AI teams don't treat evaluation as a phase — they treat it as the way they work. As the leader, you set whether quality is something the…

    Free preview
  12. Ch 12

    Build vs. Buy & Who to Hire

    You've seen what good evaluation looks like. The last decisions are yours: what to build versus buy, and who to put in the seat.

    Free preview
  13. Appendix

    Hiring an AI Evaluation Engineer

    Five questions to ask candidates, what a strong versus weak answer sounds like, and a one-page scorecard. You don't need to evaluate the AI yourself…

    Free preview

Evaluating Your AI · AI Engineer Dojo · Other editions