LIONS AISolutions
Evaluation

How to evaluate a RAG system without guessing

If you cannot say whether last week's change improved quality, you are not engineering — you are redecorating.

January 30, 2026•8 min read•Lions Ai Solutions
All insights

Start with fifty questions

You do not need a research-grade benchmark. Fifty real questions with correct answers, drawn from actual user traffic or support transcripts, catch most regressions.

Include the awkward cases deliberately: questions your documentation does not answer, questions with multiple valid answers, and questions whose answer changed recently.

Measure retrieval separately

This is the step most teams skip, and it is the most informative.

For each question, record whether the correct passage appeared in the retrieved context at all. That single number — retrieval hit rate — is a hard ceiling on end-to-end accuracy. If it is 70%, your system cannot exceed 70% no matter which model you use.

Separating retrieval from generation tells you immediately which half to work on.

Then measure the answer

Three things matter:

  • Faithfulness. Is every claim supported by the retrieved context? This catches fabrication directly.
  • Completeness. Does the answer cover what was asked, or just part of it?
  • Refusal accuracy. When the knowledge base genuinely lacks the answer, does the system say so? Test this explicitly with unanswerable questions — a system that never refuses will fabricate under pressure.

Automate it

Run the evaluation on every change to prompts, chunking, retrieval parameters or models. Wire it into CI and fail the build on regression beyond a threshold.

LLM-as-judge is adequate for faithfulness and completeness if you validate the judge against human ratings once. Retrieval hit rate needs no judge at all.

Watch it in production

Offline evaluation misses distribution shift. Track retrieval scores, refusal rates, escalation rates and latency on live traffic. A rising refusal rate usually means your content went stale, not that the model got worse.

Let's talk about what you're building

Tell us the problem. We'll tell you honestly whether AI is the right tool, and what it would take.