LIONS AISolutions
RAG

Hallucination is a retrieval problem before it is a model problem

Teams reach for a bigger model when their AI invents facts. In most systems we audit, the correct passage never reached the model at all.

April 18, 2026•7 min read•Lions Ai Solutions
All insights

The usual diagnosis is wrong

When an assistant invents a policy or a price, the instinct is to blame the model and upgrade it. We have audited enough of these systems to say that the upgrade rarely helps, because the model was never given the right information in the first place.

The question worth asking is narrow: in the cases where the system answered incorrectly, did the correct passage appear in the retrieved context?

In one telecom support system we rebuilt, the answer was no in 69% of failures. No model, at any price, answers correctly from context that does not contain the answer.

Where retrieval actually breaks

Four causes account for most of what we see.

  • Fixed-size chunking that destroys structure. Splitting every 1,000 characters cuts tables in half and separates headings from the text beneath them. A chunk reading "Unlimited after 40GB" is useless when the plan name sat in a heading three chunks earlier.
  • Pure vector search on exact terms. Embeddings are poor at rare identifiers. A query for plan code "NW-450-B" will happily match semantically similar but entirely wrong plans. Lexical search finds it instantly.
  • Follow-up questions that lose their subject. "What about the family version?" embeds to nothing useful on its own. Without rewriting against conversation history, retrieval fails.
  • Too much context. Twenty passages of marginal relevance bury the two good ones. More retrieved context past a point actively reduces accuracy.

What fixes it

Structure-aware chunking first — respect headings, keep tables intact, carry the heading breadcrumb into the embedded text.

Then hybrid retrieval. Run vector and full-text search in parallel and fuse the rankings with reciprocal rank fusion. Semantic recall and lexical precision cover each other's failure modes.

Then rerank. A second-stage relevance pass over twenty candidates, keeping the best five, consistently outperforms retrieving five directly.

Finally, instruct the model to refuse. A system that says "I don't have that information" is more valuable than one that approximates, especially where a wrong answer becomes a complaint.

Measure before you spend

Build a golden dataset — fifty real questions with correct answers is enough to start. Score retrieval separately from generation. If your retrieval hit rate is 60%, no amount of model spend will take end-to-end accuracy above it.

Let's talk about what you're building

Tell us the problem. We'll tell you honestly whether AI is the right tool, and what it would take.