AI Engineer Dojo Contents
Chapter 11

Building a RAG Culture

The teams that ship trustworthy AI over documents don't have a secret model. They have habits — a few rituals that make quality visible and keep it from sliding. Your job is to fund and protect those habits.

Culture here means a handful of concrete practices. Quality is reported as a sliced number with a sample size, not a demo. There's a gold set that a named person owns and refreshes. Bad answers from real users are mined into new test cases rather than waved away as edge cases. Releases are gated on a quality check, and a dashboard is reviewed regularly — ideally in a standing meeting, so quality is a shared number, not one engineer's private worry.

As a leader you set the incentives. If you reward flashy demos, you get flashy demos. If you consistently ask "on how many real questions, sliced how, and how do we know it didn't regress?", the team builds the muscle to answer — and the quality follows. The habits are cheap; not having them is what's expensive.

What good looks like Quality is a shared, reviewed number; the gold set is owned and growing; complaints become test cases; releases are gated; the dashboard is on the wall. Improving the eval is treated as real work, not overhead.
Red flags Quality lives in one engineer's head. No shared dashboard. Complaints dismissed as edge cases. Demos rewarded over measurements. "We'll set up evals when we have time."
Decision Lab

What you reward

In reviews, one engineer shows a slick new demo and another shows a boring chart proving last quarter's regressions are gone. Who do you praise, and what signal does it send?

How to think about it

Praise the boring chart — visibly. The demo is a snapshot that predicts nothing; the chart is evidence the product is reliably good and getting better, which is what customers actually pay for. What you celebrate in the room becomes what the team optimizes. Reward measurement and you get a team that can tell you the truth about quality; reward demos and you get a team that's great at the sixty seconds before a launch and blind to the six months after. Set the incentive deliberately.

Case study

The team that made eval a standup ritual

A product team kept shipping RAG changes that fixed one thing and quietly broke another, because quality lived in scattered heads. Their manager made one change: a single shared dashboard — retrieval quality and answer faithfulness, sliced by segment — reviewed for five minutes at the start of every standup. Suddenly regressions were everyone's business the day they appeared, not a surprise in the next customer escalation.

Within a quarter the pattern of "fixed A, broke B" faded, because no change could silently move the shared number without someone noticing. Nothing about the technology changed; the manager had simply made quality visible and collective. The habit — five minutes a day on a shared number — did more for reliability than any model upgrade they'd tried.

Running case · Brightline

Brightline puts its RAG dashboard on the wall and spends the first five minutes of each weekly standup on it. Quality stops being the eng lead's private anxiety and becomes a number the whole team owns — and it stops regressing, because now everyone would see it if it did.

Quiz · Chapter 11

  1. A healthy RAG culture treats quality as:
  2. Real user complaints should be:
  3. What a leader rewards in reviews:
  4. "We'll set up evals when we have time" is:
← Back Continue →

Chat With Your Docs · AI Engineer Dojo · aiengineerdojo.com