RAG in production is mostly a data problem
Most bad answers from a RAG system start before the model is even called. Where they come from and how to measure them.
Retrieval augmented generation, or RAG, is simple to describe: find the right pieces of your documents, give them to the model, and let it answer from them. Building a demo takes an afternoon. Making it reliable takes much longer, and most of that work is data engineering.
Two kinds of wrong answers
When a RAG system is wrong, it is usually one of two things. Either the answer contradicts the documents it was given, or it adds facts that are not in the documents at all. The fixes are different, so it helps to know which one you are looking at.
Where it really goes wrong
- Stale or duplicate documents. Three versions of the same policy and the model picks the old one.
- Bad chunking. A table split in half, or a heading separated from its content.
- Missing metadata. Without dates, owners and document types you cannot filter or rank well.
- Permissions. The search must respect who is asking. Filtering after the answer is too late.
- Retrieval that misses. If the right passage is not retrieved, no model can answer correctly.
Notice that none of these are model problems. They are pipeline, quality and ownership problems, the same ones data teams have solved for years.
Let it say "I don't know"
A good RAG system treats "the answer is not in the documents" as a correct result. Tell the model it may abstain, and reward that in your tests. A confident wrong answer costs far more trust than an honest "I could not find this."
Cite and check
Ask for a source on every claim, then check that the cited passage actually supports it. You can do this with a second, separate model call that only compares claim and passage. Doing the check in a fresh context matters, otherwise the model tends to repeat its first mistake.
What to measure
- Retrieval hit rate: how often the right passage is in the top results.
- Faithfulness: share of claims supported by the retrieved text.
- Citation accuracy: do the sources really say what the answer claims?
- Abstain rate on questions that have no answer in your documents.
Start with 50 to 100 real questions from real users, with the expected answer and source. That small set will teach you more than any benchmark.
One more rule: do not fine tune company facts into a model. Facts change. Retrieve them.