RAG and knowledge

How to evaluate a RAG system: the four numbers that tell you what broke

A wrong RAG answer has two possible causes, bad retrieval or bad generation. Measure them separately and you know which one to fix.

Cover: How to evaluate a RAG system

Most RAG projects are judged by someone typing a few questions into the chat window and deciding whether the answers feel right. That works for a demo. It fails the first time you change the chunk size, swap the embedding model, or upgrade the chat model, because you have no way to tell whether the change helped. You end up with opinions where you needed a number.

The fix is not complicated. A RAG answer can go wrong at exactly two points, and each point has metrics that look at it directly. Measure both and a bad answer stops being a mystery.

The two places a RAG answer breaks

Every RAG system does two jobs in sequence. First it retrieves passages it thinks are relevant. Then a model generates an answer from those passages.

So a wrong answer means one of these happened:

  • Retrieval missed. The passage containing the answer never reached the model. No model, however good, can answer from text it was not given.
  • Generation drifted. The right passage was there, and the model ignored it, misread it, or added claims of its own.

These need opposite fixes. A retrieval miss calls for better parsing, chunking, embeddings or search. A generation problem calls for a better prompt, a larger context window, or a stronger model. Teams that measure only the final answer tend to fix the wrong half, which is why an upgrade to a more expensive model so often changes nothing.

Two by two grid for diagnosing RAG failures. Rows: retrieval found the evidence or missed it. Columns: answer stayed faithful to the context or did not. Each cell names the metric that exposes it and the fix to try.
Retrieval metrics and generation metrics point at different fixes. Reading them together tells you which cell you are in.

Four metrics, two for each half

The names below follow Ragas, the open-source evaluation library most people meet first. Other tools use slightly different names for the same ideas.

Context recall: did retrieval find the evidence?

Context recall asks whether the retrieved passages contain what was needed to answer. In the LLM-based version, Ragas breaks a reference answer into individual claims and checks how many of them are supported by the retrieved context:

context recall = reference claims supported by retrieved context
                 / total claims in the reference

This one needs a reference answer written by a person, which is the main cost of doing evaluation properly. Ragas also offers a variant that compares chunk IDs directly, if you know which chunks should have been retrieved. That version needs no model at all.

Low recall is the clearest signal in RAG evaluation. It means the answer was never available, and nothing downstream can fix that.

Context precision: did the good passages come first?

Ragas defines context precision as the retriever’s ability to rank relevant chunks higher than irrelevant ones. Retrieving the right passage in position eight, behind seven distractors, is technically a hit and practically a problem: it wastes context window, and models pay more attention to some positions than others.

Ragas has variants with a reference answer, without one (judging against the generated response instead), and a non-LLM version based on string similarity. Precision is what you watch when you tune how many chunks to retrieve or add a re-ranker.

Faithfulness: did the answer stay inside the evidence?

Faithfulness is the generation-side metric that matters most. Ragas splits the response into claims and counts how many are supported by the retrieved context:

faithfulness = claims in the response supported by the context
               / total claims in the response

It needs no reference answer, only the question, the response and the retrieved passages, so it is cheap to run on real traffic. A low score means the model is adding things. Note what it does not tell you: an answer can be perfectly faithful to passages that were the wrong passages. That is why it has to be read next to recall.

Response relevancy: did it answer the question asked?

An answer can be faithful and still beside the point, such as a correct summary of a policy when the user asked about one exception inside it. Response relevancy scores how well the response addresses the question itself. It catches evasive, padded and off-topic answers that the other three metrics let through.

Build the test set before you pick a tool

Every metric above is only as good as the questions you run it on. The test set is the real work, and it is worth more than the choice of library.

  1. Collect real questions. Pull them from support tickets, search logs, or people who will use the system. Invented questions tend to be easier than real ones.
  2. Write the reference answer for each, in a sentence or two, and note which document and section it comes from. This is what context recall needs.
  3. Include hard cases on purpose: questions whose answer is in a table, questions that need two documents, and at least a few questions whose honest answer is “that is not in the documents.”
  4. Freeze it. Keep the set fixed while you tune. Add new questions in a separate batch so old scores stay comparable.

Fifty well-chosen questions will tell you more than five hundred generated ones. Start small and make each question earn its place.

Retrieval metrics you can compute without any model

If your test set records which chunks should be retrieved, you can measure the retrieval half with plain code: no evaluation library, no LLM judge, no cost per run. This makes it practical to rerun on every change.

def recall_at_k(retrieved_ids, relevant_ids, k):
    """Share of the relevant chunks that appear in the top k results."""
    if not relevant_ids:
        return None
    top = set(retrieved_ids[:k])
    return len(top & set(relevant_ids)) / len(relevant_ids)


def reciprocal_rank(retrieved_ids, relevant_ids):
    """1 / position of the first relevant chunk; 0 if none was retrieved."""
    relevant = set(relevant_ids)
    for position, chunk_id in enumerate(retrieved_ids, start=1):
        if chunk_id in relevant:
            return 1 / position
    return 0.0


def evaluate(test_set, search, k=5):
    recalls, ranks = [], []
    for case in test_set:
        got = search(case["question"])          # list of chunk IDs, best first
        r = recall_at_k(got, case["relevant_ids"], k)
        if r is not None:
            recalls.append(r)
        ranks.append(reciprocal_rank(got, case["relevant_ids"]))
    return {
        f"recall@{k}": sum(recalls) / len(recalls),
        "mrr": sum(ranks) / len(ranks),
    }

Recall@k is the ID-based cousin of context recall. Mean reciprocal rank rewards getting a relevant chunk to the top, which is what context precision cares about. One practical catch: chunk IDs change when you re-chunk documents. Store the relevant passage as text, and map it to new IDs after each re-index, or your scores will drop for reasons that have nothing to do with quality.

A few things LLM-judged metrics get wrong

Faithfulness and relevancy are scored by a model, and a model grading a model has limits worth keeping in mind:

  • Scores move when the judge changes. Pin the judge model and its version, and do not compare numbers produced by different judges.
  • Averages hide the pattern. A faithfulness of 0.9 across a test set can mean every answer is slightly loose or one in ten is badly wrong. Look at the worst cases individually.
  • Spot-check the judge. Read twenty of its verdicts yourself. If you disagree with several, fix that before trusting any dashboard built on it.

Agentic RAG adds a third place to break

When retrieval is done by an agent that decides what to search for, and when, a new failure appears: the agent can pick the wrong tool, search with the wrong query, or stop too early. Ragas ships separate metrics for this, including tool call accuracy, agent goal accuracy and topic adherence. The four metrics above still apply to each retrieval step, but they will not tell you the agent never searched at all. If you are deciding whether you need an agent for retrieval in the first place, AI agents vs workflows is a good place to start.

Making the numbers useful

Run the full set before and after every meaningful change, and record the scores next to what changed. Within a few rounds you will have something no demo gives you: evidence for which settings matter for your documents. That also makes tool comparisons honest. If you are choosing between a light setup such as AnythingLLM with Ollama and a heavier engine such as RAGFlow, run the same test set through both and let recall on your hard questions decide.

Comments

No comments yet — be the first to share what you think.

Leave a comment

Your email address stays private. Required fields are marked