LIONS AISolutions
Telecom • Northwind Telecom

Cutting support volume by half with grounded retrieval

A national operator's support assistant answered 19% of queries correctly. Rebuilding the retrieval layer took it to 54% without changing the model.

54%
Contacts deflected (from 19%)
-38%
Average handle time
92%
Retrieval hit rate
-61%
Inference cost per conversation
The challenge

Where they started

Northwind had deployed an off-the-shelf chatbot over 4,000 pages of tariff and policy documentation. It deflected 19% of contacts and generated billing complaints when it guessed. The vendor's recommendation was a larger, more expensive model.

What we did

  1. 1

    Built a 400-question golden dataset from real contact-centre transcripts and scored the existing system against it to establish a baseline.

  2. 2

    Diagnosed the failure as retrieval, not generation: the correct passage reached the model in only 31% of failing cases.

  3. 3

    Replaced naive fixed-size chunking with structure-aware chunking that preserved tariff table hierarchy.

  4. 4

    Introduced hybrid retrieval — vector similarity fused with full-text search — so exact plan codes and product names matched reliably.

  5. 5

    Added a reranking pass and a strict refusal instruction for low-confidence retrieval.

  6. 6

    Wired automatic escalation to a human agent whenever retrieval confidence fell below threshold.

The outcome

Correct-answer rate rose from 19% to 54% on the same model. Billing-related complaints attributable to the assistant fell to near zero because the system now refuses rather than guesses.

“The difference was grounding, not a bigger model. We were about to spend four times as much on inference to fix a chunking problem.”
— Head of Customer Operations, Northwind Telecom

Let's talk about what you're building

Tell us the problem. We'll tell you honestly whether AI is the right tool, and what it would take.