Start with fifty questions
You do not need a research-grade benchmark. Fifty real questions with correct answers, drawn from actual user traffic or support transcripts, catch most regressions.
Include the awkward cases deliberately: questions your documentation does not answer, questions with multiple valid answers, and questions whose answer changed recently.
Measure retrieval separately
This is the step most teams skip, and it is the most informative.
For each question, record whether the correct passage appeared in the retrieved context at all. That single number — retrieval hit rate — is a hard ceiling on end-to-end accuracy. If it is 70%, your system cannot exceed 70% no matter which model you use.
Separating retrieval from generation tells you immediately which half to work on.
Then measure the answer
Three things matter:
- Faithfulness. Is every claim supported by the retrieved context? This catches fabrication directly.
- Completeness. Does the answer cover what was asked, or just part of it?
- Refusal accuracy. When the knowledge base genuinely lacks the answer, does the system say so? Test this explicitly with unanswerable questions — a system that never refuses will fabricate under pressure.
Automate it
Run the evaluation on every change to prompts, chunking, retrieval parameters or models. Wire it into CI and fail the build on regression beyond a threshold.
LLM-as-judge is adequate for faithfulness and completeness if you validate the judge against human ratings once. Retrieval hit rate needs no judge at all.
Watch it in production
Offline evaluation misses distribution shift. Track retrieval scores, refusal rates, escalation rates and latency on live traffic. A rising refusal rate usually means your content went stale, not that the model got worse.
