Two different failure modes
Vector search understands meaning. Ask for "cheapest plan for someone who travels a lot" and it finds the roaming tariff page even though none of those words appear on it.
Full-text search understands tokens. Ask for "NW-450-B" and it finds exactly that string.
Each fails where the other succeeds. Vector search degrades badly on rare identifiers, part numbers, surnames and acronyms — these are poorly represented in embedding space and semantically similar strings are often entirely wrong answers. Lexical search fails whenever the user's vocabulary differs from the document's.
Production queries contain both kinds of language, frequently in the same sentence.
Fusing the rankings
The practical approach is reciprocal rank fusion. Run both searches, then score each document by position rather than by raw score:
score(d) = Σ weight_i / (k + rank_i(d))with k around 60. Documents ranked highly by either retriever surface; documents ranked well by both surface strongly.
RRF's advantage is that it never compares a cosine similarity to a BM25 score directly. Those numbers are not commensurable and normalising them is fragile. Ranks are.
Weighting
Start at equal weight and tune against your evaluation set. We typically end up weighting vector slightly higher for conversational products and lexical higher for technical documentation dense with identifiers.
Expose the weights as configuration. The right balance shifts as a corpus grows, and you want to change it without a deployment.
One caveat
Hybrid retrieval doubles your query cost and adds latency. For a corpus of a few hundred well-written documents with consistent vocabulary, vector search alone may be sufficient. Measure before adding machinery.
