RAG

RAG Architecture Deep Dive: Vector Embeddings, Hybrid Search & Reranking

A production engineering guide to building accurate Retrieval-Augmented Generation pipelines using dense embeddings, BM25 keyword matching, and cross-encoder rerankers.

Diagram for RAG Architecture Deep Dive: Vector Embeddings, Hybrid Search & Reranking
On this page

Retrieval-Augmented Generation (RAG) bridges private knowledge bases with the generative capabilities of large language models. However, standard naive RAG pipelines suffer from low precision, irrelevant context retrieval, and hallucinations.

In this guide, we explore the multi-stage architecture required for enterprise-grade retrieval.

The Limitations of Naive RAG

Naive RAG relies solely on cosine similarity over fixed-chunk vector embeddings:

  1. Text is split into naive 500-token chunks.
  2. An embedding model converts chunks into dense vectors.
  3. User questions are embedded and top-$k$ nearest neighbors are retrieved.

This fails whenever the user's query depends on exact product codes, acronyms, or multi-hop logic that dense semantic embeddings tend to blur.

Stage 1: Chunking with Context Preservation

Rather than slicing text arbitrarily at character counts, effective chunking respects markdown structure and document semantics:

python
from typing import List
 
def chunk_markdown_by_section(markdown_text: str) -> List[dict]:
    """Splits markdown by headings to maintain contextual coherence."""
    sections = markdown_text.split("\n## ")
    chunks = []
    
    for idx, section in enumerate(sections):
        if not section.strip():
            continue
        lines = section.split("\n")
        title = lines[0].replace("#", "").strip()
        body = "\n".join(lines[1:]).strip()
        
        chunks.append({
            "chunk_id": f"chunk-{idx}",
            "header": title,
            "text": f"Section: {title}\n{body}"
        })
    return chunks

Stage 2: Hybrid Search with Reciprocal Rank Fusion

Hybrid search combines the semantic generalization of dense embeddings with the exact keyword precision of sparse BM25:

python
def reciprocal_rank_fusion(dense_ranks: list, sparse_ranks: list, k: int = 60) -> list:
    """Combines ranking lists using the RRF algorithm."""
    scores = {}
 
    for rank, doc_id in enumerate(dense_ranks):
        scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
 
    for rank, doc_id in enumerate(sparse_ranks):
        scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank + 1)
 
    # Sort documents by combined fusion score
    sorted_docs = sorted(scores.items(), key=lambda item: item[1], reverse=True)
    return [doc_id for doc_id, score in sorted_docs]

Stage 3: Cross-Encoder Reranking

Bi-encoder embedding models compute query and document representations independently. Cross-encoders, on the other hand, evaluate both simultaneously through all self-attention layers:

python
from sentence_transformers import CrossEncoder
 
reranker = CrossEncoder('cross-encoder/ms-marco-MiniLM-L-6-v2')
 
def rerank_top_k(query: str, retrieved_docs: list, top_n: int = 3) -> list:
    pairs = [[query, doc["text"]] for doc in retrieved_docs]
    scores = reranker.predict(pairs)
    
    for doc, score in zip(retrieved_docs, scores):
        doc["score"] = float(score)
        
    return sorted(retrieved_docs, key=lambda d: d["score"], reverse=True)[:top_n]

Conclusion

A modern RAG architecture is not a single vector lookup. It is an optimized multi-tier information retrieval system that balances recall at the candidate generation layer with high precision at the reranking layer.

Keep learning with Sri

More practical tutorials and experiments on the channel.

Watch on YouTube
Back to articles