- Hybrid search runs keyword search (such as BM25) and vector search on the same query, then merges the two ranked lists.
- Vector search finds passages with similar meaning; keyword search finds exact terms like IDs, error codes, names and rare jargon that embeddings blur.
- Reciprocal rank fusion (RRF) merges lists by rank instead of raw score, so the two scoring systems never need to be calibrated against each other.
- A cross-encoder re-ranker over the top 20 to 50 merged results usually adds more precision than any change to the embedding model.
Vector search made retrieval feel solved. Embed the documents, embed the question, return the nearest neighbours. Then a user asks about error E4012, or order ORD-1002, or a product called “Atlas Pro”, and the top results are about something else entirely.
01Why does pure vector search miss exact matches?
An embedding compresses a passage into a point in a semantic space. That is exactly what makes it good at paraphrase: “cancel my subscription” lands near “how do I stop being billed”. It is also why it is weak at exact strings.
Identifiers, codes, version numbers and rare names carry little semantic signal. To an embedding model, E4012 and E4021 look nearly identical, and a product name it has never seen gets mapped near generic words. The passage that contains the exact token may rank below passages that are merely about errors in general.
Keyword search has the opposite profile. BM25 scores documents by how often the query terms appear, weighted so that rare terms count more. It finds E4012 instantly and has no idea that “stop being billed” means “cancel”.
02What is hybrid search?
Run both searches on the same query and merge the results. Each covers the other's blind spot.
def hybrid_search(query: str, scope: Scope, k: int = 8) -> list[Chunk]:
keyword = bm25_index.search(query, filter=scope, top_k=50)
semantic = vector_index.search(embed(query), filter=scope, top_k=50)
fused = reciprocal_rank_fusion([keyword, semantic], k=60)
return rerank(query, fused[:40])[:k]
def reciprocal_rank_fusion(result_lists, k=60):
scores = defaultdict(float)
for results in result_lists:
for rank, chunk in enumerate(results, start=1):
scores[chunk.id] += 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)03Why merge by rank instead of score?
The obvious approach is to normalise both scores and add them. It is fragile. BM25 scores are unbounded and depend on the corpus; cosine similarities sit in a narrow band that shifts by embedding model. Any weighting you tune breaks when the corpus or the model changes.
Reciprocal rank fusion (RRF) ignores raw scores. A document's fused score is the sum of 1 / (k + rank) across the lists it appears in, with k usually set around 60. A document ranked highly by both searches rises to the top; a document found by only one still gets in, just lower. There is nothing to calibrate.
04Where does re-ranking fit?
Both first-stage searches are built for speed over millions of chunks, so they compare the query and the document separately. A cross-encoder re-ranker reads the query and each candidate passage together and scores how well the passage actually answers the question. It is far more accurate and far too slow to run on the whole corpus.
So use it as a second stage: take the top 20 to 50 fused results and re-rank them. That small, accurate step is often the largest single precision gain in a retrieval pipeline.
Embeddings know what you mean. Keywords know what you said. Production queries need both.
05Is it worth the extra infrastructure?
Usually yes, and it is less extra than it sounds. Many search engines and vector databases now support keyword and vector search in one index, and BM25 over a few million chunks is cheap. The permission filter applies to both searches in the same way, inside the query (see filter by permission before you rank).
Measure before and after on the same question set, as described in chunk by document structure: recall at k for retrieval, and answer accuracy end to end. Questions with identifiers are where the difference shows first.