← Engineering notes

Note 01 · RAG · Evaluation

How we evaluate a RAG system before it goes live.

We never ship a RAG system without running it through our eval harness first. Seven metrics, three ground-truth sets, a regression report on every model or index change. This note describes what we measure, how we build the datasets and why the usual demo answers are misleading.

~18 min May 2026 From three production deployments
TL;DR

Ein RAG-System auf Basis von „Demo läuft” zu shippen ist verantwortungslos. Wir messen retrieval-Qualität (recall@5, MRR, hit-rate@k) getrennt von generation-Qualität (faithfulness, groundedness, answer relevance) und tracken beides gegen ein versioniertes Ground-Truth-Set aus echten Kundenfragen. Jeder Commit, der Embeddings, Chunk-Strategie oder den Generator anfasst, läuft automatisch durch das Harness. Regression > 5% blockt den Merge.

1 · Why the usual demo lies

Every RAG demo asks the same three questions the developer used while building. They work — because the system has been indirectly optimized for those exact questions. As soon as a real clerk, lawyer or clinic staffer phrases things differently, recall and faithfulness drop.

We learned this in an early project: an internal knowledge bot looked brilliant in the demo, but came back to us two weeks after launch with 31% incorrect answers. Since then, the evaluation harness has been a mandatory part of our process.

2 · The seven metrics

We strictly separate “does the system find the relevant documents?” (retrieval) from “does the system answer truthfully based on those documents?” (generation). Both stages can break independently.

Stage Metrik What it measures Threshold
Retrievalrecall@5Share of questions where the right document is in the top 5≥ 0.92
MRR@10Mean Reciprocal Rank — how high the right doc ranks≥ 0.78
hit-rate@1Top-1 is the right doc (for strict-mode answers)≥ 0.65
GenerationfaithfulnessIs the answer fully supported by the retrieved context?≥ 0.88
answer-relevanceDoes the system actually answer the question asked?≥ 0.85
refusal-rateShare of correct refusals on unanswerable questions≥ 0.90
Systemlatency P95End-to-end latency from request to finished answer≤ 2.4 s

Thresholds are not universal — we calibrate them per use case. A legal research system has a much higher refusal-rate requirement than a marketing assistant. What matters is that the numbers are signed off by the customer before the first sprint.

3 · Building the ground-truth set

This is the unglamorous, time-consuming part — and the most important. Per project we build three separate sets:

  • Set A — reality (80–120 questions): from customer tickets, emails, call notes. With realistic phrasings, typos, half-sentences. Per question: the canonical answer plus the IDs of the documents that support it.
  • Set B — adversarial (30–50 questions): deliberately misleading or ambiguous queries, questions about information not in the corpus, out-of-scope requests. Here the refusal rate matters most.
  • Set C — regression (grows over time): every question that was once answered wrong in production lands here. With the correct answer. That way we see immediately when a future model swap reintroduces old bugs.

We version these sets in a private repo next to the code. When a customer ships us a new document that changes answers, we update the set. It's work. It pays off every time.

4 · Die Pipeline

The harness is a modest Python script that runs against the production (or staging) pipeline. No custom framework, no “RAG Eval Studio” — we want readable code that can be changed in 20 minutes.

$ uv run python eval/run.py --set A --tag pre-deploy-2026-05

[1/3] Loading ground-truth set A ............... 104 questions
[2/3] Running pipeline (vllm@l40s, top-k=5) ... ████████ 104/104
[3/3] Computing metrics ........................ done

retrieval/recall@5 ............ 0.943   (Δ +0.012 vs baseline)
retrieval/MRR@10 .............. 0.812   (Δ -0.004 vs baseline)
retrieval/hit-rate@1 .......... 0.683   (Δ +0.021 vs baseline)
generation/faithfulness ....... 0.891   (Δ -0.018 vs baseline)
generation/answer-relevance ... 0.876   (Δ +0.005 vs baseline)
generation/refusal-rate ....... 0.923   (Δ -0.011 vs baseline)
system/latency-p95 ............ 1.94 s  (Δ -0.21 s  vs baseline)

✓ all thresholds met. Δ-faithfulness within ±0.02 tolerance.
Report written to: eval/reports/2026-05-27-pre-deploy.html

An HTML report file lands in a directory tracked by the repo. On every PR that touches generator or retriever, the harness runs automatically — no merge without a green check.

5 · What we're probably doing wrong

Ehrlich, weil das nicht gelöst ist:

  • Faithfulness mit LLM-as-Judge gemessen ist noisy. Wir verwenden zwei verschiedene Judges (lokales Llama 3.1 + Claude als Zweit-Judge) und nehmen den niedrigeren Score. Wichtig: Der Zweit-Judge läuft ausschließlich auf anonymisierten bzw. synthetischen Eval-Sets — es gehen keine Kundendaten an US-Anbieter. Trotzdem schwanken Werte ±0.03 zwischen Läufen.
  • Set A is never large enough. One hundred questions do not cover the long tail. We compensate with set C, but new domains leave blind spots in the first weeks.
  • Latenz auf Staging-Hardware ist optimistisch. Wir haben gelernt, P95 immer +30% draufzuschlagen, bevor wir versprochene SLAs gegenüber dem Kunden formulieren.

6 · Why this is decisive for regulated industries

Eine Hörgeräte-Klinik kann nicht 31% falsche Auskünfte tolerieren. Ein Anwalt nicht 5% halluzinierte Paragrafen. Eine Industrie-Wartung nicht 10% irreführende Reparaturanweisungen. Für genau diese Kunden bauen wir RAG. Das Eval-Harness ist der einzige Weg, vor dem ersten Live-Anruf zu wissen, ob das System die Anforderung trifft.

If a vendor sells you a RAG system and says nothing about eval, ground-truth or regression — ask. If the answer stays vague, change vendors.