The RAG Paradigm: Why Retrieval Matters for LLMs
Large language models carry a fundamental constraint: their knowledge is frozen at training time. Any fact, policy document, or internal dataset created after the training cutoff is invisible to the model. For enterprise applications (compliance Q&A, internal knowledge bases, domain-specific report generation) this limitation is disqualifying on its own.
The naive solution is to concatenate all relevant documents into the prompt context. This fails for three reasons:
- Token limits are real. Even models with 128K or 200K token windows cannot ingest entire document repositories. A single corporate knowledge base can exceed millions of tokens.
- Cost scales linearly with input tokens. Feeding 100K tokens per query across thousands of daily requests produces unsustainable API bills.
- Accuracy degrades with context length. Benchmarks on the BABILong evaluation show GPT-4 Turbo dropping from 100% accuracy at 1K context to 6% at 128K context on single-fact QA tasks. GPT-4o shows comparable but less severe degradation, falling from 100% to 27% over the same range. The model has the information but fails to locate it.
Retrieval-Augmented Generation (RAG) addresses all three problems. Rather than stuffing the full corpus into the prompt, RAG retrieves only the document fragments most relevant to a given query and injects those fragments as context before generation. The LLM API then synthesizes an answer grounded in retrieved evidence, with the option to cite sources.
RAG was introduced by Lewis et al. (2020) as a method to combine parametric memory (the model weights) with non-parametric memory (an external retrieval index). The approach has since become the standard architecture for grounding LLM outputs in enterprise data.
When to use RAG:
- You need answers grounded in proprietary or frequently updated data
- Reducing hallucinations is critical (compliance, legal, medical contexts)
- Queries require verifiable, source-attributed responses
When RAG is not the right fit:
- Response latency must be minimal and real-time retrieval is a bottleneck
- The dataset is small, static, and fits entirely within the model's context window
- The task requires end-to-end creative generation without grounding constraints
Trade-offs to consider:
| Pros | Cons |
|---|---|
| Reduces hallucinations and errors | Adds latency due to the retrieval step |
| Enables use of fresh, proprietary data | Increases system architecture complexity |
| Transparent with source citations | Adds cost for embedding storage and queries |
| Modular: update knowledge base without retraining | Requires robust chunking and retrieval strategies |
Architecture Overview: The Indexing and Retrieval Pipeline
A RAG system operates in two distinct phases: an offline indexing pipeline and an online retrieval-generation pipeline.
Indexing Pipeline (Offline)
The indexing pipeline prepares your document corpus for efficient semantic search:
-
Document Loading. Ingest raw data from sources: PDFs, databases, APIs, web scrapers, Confluence pages, Slack exports. Each source requires a dedicated loader that extracts clean text and preserves structural metadata (titles, headers, page numbers).
-
Chunking. Split documents into smaller text segments. Chunk size directly affects retrieval precision. This step is covered in detail below.
-
Embedding. Pass each chunk through an embedding model to produce a dense vector representation. The embedding captures semantic meaning, not just keyword overlap.
-
Storage. Write the vectors, along with the original text and metadata, into a vector database. The database builds an index (typically HNSW or IVF) for approximate nearest neighbor search.
Retrieval-Generation Pipeline (Online)
At query time, the system executes four steps:
- Query Embedding. The user's question is embedded using the same model that encoded the document chunks.
- Retrieval. The vector database performs a similarity search (cosine similarity or dot product) to return the top-K most relevant chunks.
- Augmentation. Retrieved chunks are formatted and injected into a prompt template alongside the user's question.
- Generation. The LLM produces a response conditioned on both the question and the retrieved context.
The critical architectural insight is that retrieval and generation are decoupled. You can swap embedding models, change vector databases, adjust chunk sizes, or add re-ranking layers without modifying the generation component. Good prompting strategies for the generation step further improve answer quality by instructing the LLM on how to use the retrieved context.
Embeddings: How Text Becomes Vectors
An embedding model maps a text string to a fixed-dimensional vector in a continuous space where semantic similarity corresponds to geometric proximity. Two sentences about the same topic will have vectors with high cosine similarity, even if they share no common words.
Embedding Model Selection
The Massive Text Embedding Benchmark (MTEB) provides standardized evaluation across 8 task categories and 58 datasets, covering clustering, retrieval, classification, semantic textual similarity, and re-ranking. Key models to consider:
| Model | Dimensions | Context Window | Strengths |
|---|---|---|---|
OpenAI text-embedding-3-large | 3072 (configurable) | 8191 tokens | Strong English performance, dimension reduction via Matryoshka |
Cohere embed-v3 | 1024 | 512 tokens | Multilingual, built-in compression types |
Voyage voyage-large-2 | 1536 | 16000 tokens | Long-context embedding, strong on code |
BGE bge-large-en-v1.5 | 1024 | 512 tokens | Open-source, self-hostable, no API dependency |
How Embeddings Capture Meaning
Word embeddings encode semantic relationships as vector arithmetic. The classic example: the vector for "king" minus "man" plus "woman" yields a vector close to "queen." Modern sentence-level embeddings extend this principle to entire passages, capturing topic, intent, and nuance in a single vector.
The embedding space clusters related concepts. Documents about financial risk will occupy a region distant from documents about recipe ingredients, regardless of shared vocabulary. This geometric structure is what makes semantic search possible. You are searching by meaning, not by keyword match.
Similarity Metrics
The choice of distance metric affects retrieval behavior:
- Cosine similarity. Measures the angle between vectors, ignoring magnitude. Standard choice for normalized embeddings. Range: [-1, 1].
- Dot product. Equivalent to cosine similarity when vectors are unit-normalized. Faster to compute. Preferred when embeddings are pre-normalized.
- Euclidean distance (L2). Measures absolute distance. Sensitive to vector magnitude. Less common in RAG pipelines.
Most embedding providers normalize their outputs, making cosine similarity and dot product interchangeable in practice.
Challenges in Embedding Models
Embedding models come with practical limitations worth accounting for during system design:
- Context sensitivity. A single word can carry different meanings depending on context. Embedding models may not always capture these nuances, leading to imprecise retrieval for ambiguous queries.
- Scalability. Large document corpora generate millions of vectors. Managing index size, memory consumption, and query latency requires careful infrastructure planning.
- Computational costs. Training and running embedding models (especially large ones) demands significant GPU resources. For self-hosted models, this translates to non-trivial infrastructure investment.
- Bias and ethical implications. Embedding models inherit biases from their training data. Results can vary across runs, and bias in the embedding space can surface in retrieval quality, particularly for underrepresented topics or languages.
Chunking Strategies: Fixed-Size, Semantic, Recursive
Chunking determines how source documents are divided into retrievable units. Chunk size and strategy directly control the trade-off between retrieval precision and context completeness.
Fixed-Size Chunking
The simplest approach: split text into segments of N tokens (or characters) with an overlap of M tokens between consecutive chunks.
from langchain_text_splitters import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
separators=["\n\n", "\n", ". ", " ", ""]
)
chunks = splitter.split_text(document_text)
The overlap parameter maintains contextual continuity at chunk boundaries. Without overlap, a sentence split across two chunks loses coherence in both. Typical overlap values range from 10-20% of chunk size.
Chunk size trade-offs:
- Small chunks (128-256 tokens): Higher retrieval precision but risk losing context needed for a complete answer. More chunks to store and search.
- Large chunks (1024-2048 tokens): More context per retrieval hit but lower precision, since irrelevant text dilutes the signal. Fewer total chunks, faster indexing.
- Common default: 512 tokens with 64-token overlap balances precision and context for most use cases.
Semantic Chunking
Instead of splitting at fixed intervals, semantic chunking respects document structure. Splits occur at paragraph boundaries, section headers, or topic transitions.
from langchain_experimental.text_splitter import SemanticChunker
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
semantic_splitter = SemanticChunker(
embeddings,
breakpoint_threshold_type="percentile",
breakpoint_threshold_amount=95
)
semantic_chunks = semantic_splitter.split_text(document_text)
Semantic chunking uses embedding similarity between consecutive sentences. When the similarity drops below a threshold (indicating a topic shift), a chunk boundary is inserted. This preserves the semantic integrity of each chunk at the cost of variable chunk sizes.
Specialized Chunking
For structured documents (HTML, Markdown, code), use format-aware splitters that respect the document hierarchy:
- Markdown headers. Split on
#,##,###boundaries, preserving header metadata - HTML tags. Split on
<section>,<article>,<div>elements - Code. Split on function/class definitions, preserving complete callable units
The choice of chunking strategy depends on your corpus. Homogeneous text (news articles, reports) works well with fixed-size chunking. Heterogeneous collections (mixed PDFs, structured documents, code repositories) benefit from specialized or semantic approaches.
Vector Databases: Pinecone, Weaviate, ChromaDB Comparison
The vector database stores embeddings and serves similarity queries at scale. The ecosystem divides into three categories:
- Purpose-built vector databases: Pinecone, Weaviate, Milvus, Qdrant, Chroma
- Vector-capable traditional databases: PostgreSQL (pgvector), MongoDB Atlas Vector Search, Redis, Neo4j
- Vector libraries (in-process): FAISS, Annoy, HNSWlib
Comparison for RAG Workloads
| Feature | Pinecone | Weaviate | ChromaDB |
|---|---|---|---|
| Deployment | Fully managed (SaaS) | Self-hosted or managed cloud | Embedded (in-process) or client-server |
| Indexing algorithm | Proprietary (HNSW-based) | HNSW | HNSW |
| Metadata filtering | Native, pre-filter | Native, GraphQL API | Native, where clauses |
| Hybrid search | Sparse-dense vectors | BM25 + vector (built-in) | Not built-in |
| Max vectors | Billions (serverless) | Millions (depends on resources) | Millions (limited by memory) |
| Best for | Production SaaS, minimal ops | Self-hosted with fine control, GraphQL ecosystem | Prototyping, local development, small-scale apps |
| Pricing model | Per-query + storage | Open-source (self-host) or managed pricing | Open-source, free |
Selection Criteria
- Prototyping and experimentation: ChromaDB. Zero infrastructure, pip-installable, works in a Jupyter notebook.
- Production with managed infrastructure: Pinecone. No operational burden, scales automatically, strong SLA guarantees.
- Production with self-hosting requirements: Weaviate or Milvus. Full control over data residency and infrastructure, with built-in hybrid search capabilities.
- Existing Postgres infrastructure: pgvector extension. Avoids introducing a new database into the stack, though it trades off vector search performance at scale.
Implementation: Building a RAG Pipeline with LangChain
The following implementation demonstrates a complete RAG pipeline using LangChain, OpenAI embeddings, and ChromaDB as the vector store. This code is production-readable but uses ChromaDB for simplicity. Swap to Pinecone or Weaviate for production scale.
Step 1: Document Loading and Chunking
from langchain_community.document_loaders import PyPDFLoader
from langchain_text_splitters import RecursiveCharacterTextSplitter
# Load documents
loader = PyPDFLoader("corporate_knowledge_base.pdf")
documents = loader.load()
# Split into chunks
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=64,
separators=["\n\n", "\n", ". ", " ", ""]
)
chunks = text_splitter.split_documents(documents)
print(f"Split {len(documents)} pages into {len(chunks)} chunks")
Step 2: Embedding and Vector Store Creation
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
embedding_model = OpenAIEmbeddings(model="text-embedding-3-small")
vectorstore = Chroma.from_documents(
documents=chunks,
embedding=embedding_model,
collection_name="knowledge_base",
persist_directory="./chroma_db"
)
retriever = vectorstore.as_retriever(
search_type="similarity",
search_kwargs={"k": 5}
)
Step 3: RAG Chain Assembly
from langchain_openai import ChatOpenAI
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
llm = ChatOpenAI(model="gpt-4o", temperature=0)
prompt = ChatPromptTemplate.from_template("""
Answer the question based only on the following context.
If the context does not contain enough information, say so explicitly.
Cite the source document and page number when possible.
Context:
{context}
Question: {question}
Answer:
""")
def format_docs(docs):
return "\n\n---\n\n".join(
f"[Source: {doc.metadata.get('source', 'unknown')}, "
f"Page: {doc.metadata.get('page', 'N/A')}]\n{doc.page_content}"
for doc in docs
)
rag_chain = (
{"context": retriever | format_docs, "question": RunnablePassthrough()}
| prompt
| llm
| StrOutputParser()
)
# Query the pipeline
response = rag_chain.invoke("What are the key risk factors mentioned in the report?")
print(response)
This three-step pipeline covers the core RAG workflow. The retriever | format_docs composition handles the retrieval and context formatting, while LangChain Expression Language (LCEL) chains the components declaratively.
RAG vs Long Context Windows: When to Use Which
Modern LLMs offer increasingly large context windows. GPT-4o supports 128K tokens, and Gemini 1.5 Pro handles up to 1M tokens. This raises a natural question: is RAG still necessary?
The answer depends on five factors:
Cost
RAG retrieves 5-20 chunks (roughly 2,500-10,000 tokens) per query. Feeding 100K tokens of raw documents per query costs 10-40x more in API fees. For high-volume applications (thousands of queries per day), RAG reduces inference cost by an order of magnitude.
Accuracy at Scale
Long-context performance degrades as the window fills. On the HotpotQA benchmark, RAG with retrieved context consistently outperforms stuffing full documents into long context windows across multiple model families. The retrieval step acts as a relevance filter, ensuring the model attends to the right information rather than searching through noise.
Latency
Processing 100K tokens takes significantly longer than processing 3K tokens of retrieved context. Time-to-first-token scales with input length. RAG adds a retrieval step (typically 50-200ms for vector search) but reduces generation latency by minimizing input size.
Data Freshness
RAG indexes can be updated incrementally. When a new document arrives, embed and insert it. No model retraining or prompt reconstruction required. Long-context approaches require rebuilding the full prompt on every query from the latest corpus.
When Long Context Wins
Long context is preferable when:
- The entire corpus fits within a single context window (small dataset, single document)
- The task requires reasoning across the entire document (summarization, holistic analysis)
- Implementation simplicity outweighs cost concerns (prototyping, low-volume use cases)
In practice, many production systems combine both approaches: RAG for retrieval from large corpora, with retrieved chunks inserted into a reasonably sized context window that still leaves room for multi-turn conversation history.
Optimization Techniques: Re-ranking, Hybrid Search, Query Expansion
A baseline RAG pipeline retrieves the top-K chunks by vector similarity and passes them directly to the LLM. Several techniques improve retrieval quality beyond this baseline.
Re-ranking
Vector search returns an initial candidate set (top-N, where N = 25-100). A cross-encoder re-ranker then scores each candidate against the original query with higher fidelity than bi-encoder similarity, producing a refined top-K (K = 3-5) for the final prompt.
from langchain.retrievers import ContextualCompressionRetriever
from langchain_cohere import CohereRerank
compressor = CohereRerank(
model="rerank-english-v3.0",
top_n=5
)
reranking_retriever = ContextualCompressionRetriever(
base_compressor=compressor,
base_retriever=vectorstore.as_retriever(search_kwargs={"k": 25})
)
# The retriever now fetches 25 candidates and re-ranks to top 5
results = reranking_retriever.invoke("What are the quarterly revenue trends?")
The two-stage architecture (broad vector recall followed by precise re-ranking) is standard in production RAG systems. Re-ranking providers include Cohere, Amazon SageMaker endpoints, and open-source cross-encoders from Hugging Face (e.g., cross-encoder/ms-marco-MiniLM-L-12-v2).
Hybrid Search
Pure vector search misses exact keyword matches. A query for document ID "DOC-2024-0847" will fail in embedding space because the string has no semantic content. Hybrid search combines dense vector retrieval with sparse keyword retrieval (BM25) to cover both semantic and lexical matching.
Weaviate supports hybrid search natively with a configurable alpha parameter that controls the weight between vector and keyword scores:
from langchain_weaviate import WeaviateVectorStore
vectorstore = WeaviateVectorStore(
client=weaviate_client,
index_name="Documents",
text_key="content",
embedding=embedding_model
)
# Hybrid search: alpha=0.75 weights 75% vector, 25% keyword
retriever = vectorstore.as_retriever(
search_type="hybrid",
search_kwargs={"alpha": 0.75, "k": 10}
)
Query Expansion
The user's query may be too terse or ambiguous for effective retrieval. Query expansion generates multiple reformulations of the original question, retrieves results for each, and merges the candidate sets.
from langchain_core.prompts import ChatPromptTemplate
from langchain_openai import ChatOpenAI
expansion_prompt = ChatPromptTemplate.from_template("""
Generate 3 alternative phrasings of this search query.
Return only the queries, one per line, no numbering.
Original query: {query}
""")
expander = expansion_prompt | ChatOpenAI(model="gpt-4o-mini", temperature=0.7)
def expanded_retrieval(query: str, retriever, k: int = 5):
import hashlib
expansions = expander.invoke({"query": query}).content.strip().split("\n")
all_queries = [query] + expansions
seen_ids = set()
unique_docs = []
for q in all_queries:
docs = retriever.invoke(q)
for doc in docs:
doc_id = hashlib.sha256(doc.page_content.encode()).hexdigest()
if doc_id not in seen_ids:
seen_ids.add(doc_id)
unique_docs.append(doc)
return unique_docs[:k]
Contextual Retrieval
A common weakness in standard chunking is that individual chunks lose their surrounding context. A chunk mentioning "the incident" may not contain enough information for the retrieval system to connect it to the right query.
Contextual retrieval addresses this by using an LLM to prepend a short context summary to each chunk before embedding. After chunking a source document, the system passes both the full document and each chunk to the LLM, asking it to generate a brief description that situates the chunk within the broader document. The contextualized chunk (summary plus original text) is then embedded and stored. This approach improves retrieval precision because each chunk now carries enough context for the embedding model to place it correctly in the vector space.
The trade-off is additional LLM cost during indexing, since every chunk requires an inference call. For large corpora, this can be managed by using smaller, cheaper models for the contextualization step.
GraphRAG: Knowledge Graphs for Structured Retrieval
Traditional vector-based RAG works well for finding semantically similar text passages, but it struggles with queries that require reasoning across relationships between entities. For example, asking about connections between people, organizations, and events requires understanding structured relationships that flat text similarity cannot capture.
GraphRAG addresses this by building a knowledge graph alongside (or instead of) a vector index. Documents are processed through an information extraction pipeline that identifies entities and their relationships, then stores them as nodes and edges in a graph database.
The retrieval process in GraphRAG combines vector similarity search with graph traversal. When a query arrives, the system finds relevant nodes through embedding similarity and then traverses the graph to discover related entities and relationships. This connected context is assembled and passed to the LLM for generation.
When GraphRAG adds value:
- Queries involve relationships between entities (e.g., organizational structures, supply chains, regulatory dependencies)
- The corpus contains highly interconnected information where context spans multiple documents
- Users need multi-hop reasoning across facts scattered in different sources
Building the knowledge graph:
Knowledge graph construction typically uses one of two approaches. If the LLM supports tool use or function calling, you define structured output schemas (Node and Relationship classes) and let the model extract entities and relations in a structured format. As a fallback, few-shot prompting can instruct the model to output extraction results as JSON.
Graph databases like Neo4j, Amazon Neptune, or open-source options like JanusGraph store the extracted graph. The resulting system can answer relationship-oriented queries that pure vector search would miss entirely.
Evaluation and Observability
Production RAG systems require continuous monitoring. Key frameworks:
- RAGAS. Open-source evaluation measuring faithfulness (does the answer match the retrieved context?), answer relevance (does it address the question?), and context precision (are retrieved documents relevant?).
- Giskard. Supports evaluation and monitoring of RAG systems with automated tests, including detection of hallucinations and retrieval failures.
- A/B testing. Systematic comparison of pipeline variants: different chunk sizes, embedding models, re-ranking configurations.
- System telemetry. Track retrieval latency, embedding latency, generation latency, cache hit rates, and cost per query. Alert on degradation.
Monitor both technical metrics and content quality metrics. Capture user interactions to identify failure patterns: frequently reformulated queries indicate retrieval gaps, while low-confidence responses suggest context insufficiency.
Multimodal RAG
Standard RAG pipelines handle text, but many enterprise knowledge bases contain images, diagrams, charts, and tables. Multimodal AI models enable RAG systems to index and retrieve across data types.
Two main approaches exist for including images in a RAG pipeline:
-
Cross-modal embeddings. A multimodal embedding model maps both text and images into a single vector space, enabling cross-modality similarity searches. A text query can retrieve relevant images, and an image query can retrieve relevant text. Models like CLIP and its successors produce aligned embeddings where semantically related text and images land near each other in the vector space.
-
Summarize-then-embed. Extract images from documents, generate text summaries of each image using a vision model, and embed the summaries alongside text chunks. This approach works with any text embedding model but adds a summarization step during indexing.
For documents containing tables, a similar strategy applies: extract the table, generate a text summary, and embed both the summary (for retrieval) and the raw table data (for the final context passed to the LLM).
Summary
RAG remains the production-standard architecture for grounding LLM outputs in external data. The core pipeline (chunk, embed, store, retrieve, generate) is straightforward, but the engineering details of each stage determine system quality. Embedding model selection, chunking strategy, vector database choice, and post-retrieval optimization (re-ranking, hybrid search, query expansion) each contribute measurably to answer accuracy and system reliability.
Start with the simplest viable pipeline: fixed-size chunking, a managed embedding API, ChromaDB for storage, and direct top-K retrieval. Measure retrieval precision and answer quality with RAGAS. Then iterate. Add re-ranking if precision is insufficient, switch to hybrid search if keyword queries fail, and adjust chunk sizes based on your document structure. For relationship-heavy domains, evaluate whether GraphRAG provides better coverage than pure vector search. The modular architecture of RAG makes each component independently tunable without rebuilding the system.