RAG in production is mostly a data problem

AI EngineeringFrom putting retrieval over company documents into production2 min read
where most answers go wrongDocsowner, dateChunksclean splitsIndex+ metadataModelanswersfix the data before you blame the model

Most bad answers from a RAG system start before the model is even called. Where they come from and how to measure them.

Retrieval augmented generation, or RAG, is simple to describe: find the right pieces of your documents, give them to the model, and let it answer from them. Building a demo takes an afternoon. Making it reliable takes much longer, and most of that work is data engineering.

Two kinds of wrong answers

When a RAG system is wrong, it is usually one of two things. Either the answer contradicts the documents it was given, or it adds facts that are not in the documents at all. The fixes are different, so it helps to know which one you are looking at.

Where it really goes wrong

DocumentsClean + chunkIndextext + metadataQuestionRetrievewith permissionsDraft answerCheck claimsagainst sourcesAnswerI don't know
A RAG pipeline. The top row decides most of the quality.
  • Stale or duplicate documents. Three versions of the same policy and the model picks the old one.
  • Bad chunking. A table split in half, or a heading separated from its content.
  • Missing metadata. Without dates, owners and document types you cannot filter or rank well.
  • Permissions. The search must respect who is asking. Filtering after the answer is too late.
  • Retrieval that misses. If the right passage is not retrieved, no model can answer correctly.

Notice that none of these are model problems. They are pipeline, quality and ownership problems, the same ones data teams have solved for years.

Let it say "I don't know"

A good RAG system treats "the answer is not in the documents" as a correct result. Tell the model it may abstain, and reward that in your tests. A confident wrong answer costs far more trust than an honest "I could not find this."

Cite and check

Ask for a source on every claim, then check that the cited passage actually supports it. You can do this with a second, separate model call that only compares claim and passage. Doing the check in a fresh context matters, otherwise the model tends to repeat its first mistake.

What to measure

  • Retrieval hit rate: how often the right passage is in the top results.
  • Faithfulness: share of claims supported by the retrieved text.
  • Citation accuracy: do the sources really say what the answer claims?
  • Abstain rate on questions that have no answer in your documents.

Start with 50 to 100 real questions from real users, with the expected answer and source. That small set will teach you more than any benchmark.

One more rule: do not fine tune company facts into a model. Facts change. Retrieve them.

Faizan Khan

Faizan Khan is an AI and data engineer in Berlin. Working on something like this? Book a 30 minute call or email hello@faizankhan.me.