Note 01 · RAG · Evaluation
How we evaluate a RAG system before it goes live.
We never ship a RAG system without running it through our eval harness first. Seven metrics, three ground-truth sets, a regression report on every model or index change. This note describes what we measure, how we build the datasets and why the usual demo answers are misleading.
Ein RAG-System auf Basis von „Demo läuft” zu shippen ist verantwortungslos. Wir messen retrieval-Qualität (recall@5, MRR, hit-rate@k) getrennt von generation-Qualität (faithfulness, groundedness, answer relevance) und tracken beides gegen ein versioniertes Ground-Truth-Set aus echten Kundenfragen. Jeder Commit, der Embeddings, Chunk-Strategie oder den Generator anfasst, läuft automatisch durch das Harness. Regression > 5% blockt den Merge.
1 · Why the usual demo lies
Every RAG demo asks the same three questions the developer used while building. They work — because the system has been indirectly optimized for those exact questions. As soon as a real clerk, lawyer or clinic staffer phrases things differently, recall and faithfulness drop.
We learned this in an early project: an internal knowledge bot looked brilliant in the demo, but came back to us two weeks after launch with 31% incorrect answers. Since then, the evaluation harness has been a mandatory part of our process.
2 · The seven metrics
We strictly separate “does the system find the relevant documents?” (retrieval) from “does the system answer truthfully based on those documents?” (generation). Both stages can break independently.
| Stage | Metrik | What it measures | Threshold |
|---|---|---|---|
| Retrieval | recall@5 | Share of questions where the right document is in the top 5 | ≥ 0.92 |
| MRR@10 | Mean Reciprocal Rank — how high the right doc ranks | ≥ 0.78 | |
| hit-rate@1 | Top-1 is the right doc (for strict-mode answers) | ≥ 0.65 | |
| Generation | faithfulness | Is the answer fully supported by the retrieved context? | ≥ 0.88 |
| answer-relevance | Does the system actually answer the question asked? | ≥ 0.85 | |
| refusal-rate | Share of correct refusals on unanswerable questions | ≥ 0.90 | |
| System | latency P95 | End-to-end latency from request to finished answer | ≤ 2.4 s |
Thresholds are not universal — we calibrate them per use case. A legal research system has a much higher refusal-rate requirement than a marketing assistant. What matters is that the numbers are signed off by the customer before the first sprint.
3 · Building the ground-truth set
This is the unglamorous, time-consuming part — and the most important. Per project we build three separate sets:
- Set A — reality (80–120 questions): from customer tickets, emails, call notes. With realistic phrasings, typos, half-sentences. Per question: the canonical answer plus the IDs of the documents that support it.
- Set B — adversarial (30–50 questions): deliberately misleading or ambiguous queries, questions about information not in the corpus, out-of-scope requests. Here the refusal rate matters most.
- Set C — regression (grows over time): every question that was once answered wrong in production lands here. With the correct answer. That way we see immediately when a future model swap reintroduces old bugs.
We version these sets in a private repo next to the code. When a customer ships us a new document that changes answers, we update the set. It's work. It pays off every time.
4 · Die Pipeline
The harness is a modest Python script that runs against the production (or staging) pipeline. No custom framework, no “RAG Eval Studio” — we want readable code that can be changed in 20 minutes.
$ uv run python eval/run.py --set A --tag pre-deploy-2026-05
[1/3] Loading ground-truth set A ............... 104 questions
[2/3] Running pipeline (vllm@l40s, top-k=5) ... ████████ 104/104
[3/3] Computing metrics ........................ done
retrieval/recall@5 ............ 0.943 (Δ +0.012 vs baseline)
retrieval/MRR@10 .............. 0.812 (Δ -0.004 vs baseline)
retrieval/hit-rate@1 .......... 0.683 (Δ +0.021 vs baseline)
generation/faithfulness ....... 0.891 (Δ -0.018 vs baseline)
generation/answer-relevance ... 0.876 (Δ +0.005 vs baseline)
generation/refusal-rate ....... 0.923 (Δ -0.011 vs baseline)
system/latency-p95 ............ 1.94 s (Δ -0.21 s vs baseline)
✓ all thresholds met. Δ-faithfulness within ±0.02 tolerance.
Report written to: eval/reports/2026-05-27-pre-deploy.html
An HTML report file lands in a directory tracked by the repo. On every PR that touches generator or retriever, the harness runs automatically — no merge without a green check.
5 · What we're probably doing wrong
Ehrlich, weil das nicht gelöst ist:
- Faithfulness mit LLM-as-Judge gemessen ist noisy. Wir verwenden zwei verschiedene Judges (lokales Llama 3.1 + Claude als Zweit-Judge) und nehmen den niedrigeren Score. Wichtig: Der Zweit-Judge läuft ausschließlich auf anonymisierten bzw. synthetischen Eval-Sets — es gehen keine Kundendaten an US-Anbieter. Trotzdem schwanken Werte ±0.03 zwischen Läufen.
- Set A is never large enough. One hundred questions do not cover the long tail. We compensate with set C, but new domains leave blind spots in the first weeks.
- Latenz auf Staging-Hardware ist optimistisch. Wir haben gelernt, P95 immer +30% draufzuschlagen, bevor wir versprochene SLAs gegenüber dem Kunden formulieren.
6 · Why this is decisive for regulated industries
Eine Hörgeräte-Klinik kann nicht 31% falsche Auskünfte tolerieren. Ein Anwalt nicht 5% halluzinierte Paragrafen. Eine Industrie-Wartung nicht 10% irreführende Reparaturanweisungen. Für genau diese Kunden bauen wir RAG. Das Eval-Harness ist der einzige Weg, vor dem ersten Live-Anruf zu wissen, ob das System die Anforderung trifft.
If a vendor sells you a RAG system and says nothing about eval, ground-truth or regression — ask. If the answer stays vague, change vendors.