- Empty retrieval results in production often come from preprocessing inconsistencies between document ingestion and query handling — Unicode normalization and caching deterministic preprocessing fixes 40% of zero-result queries.
- Chunk boundaries destroy context when answers span multiple chunks — smart chunking with overlap and boundary metadata lets you fetch adjacent chunks automatically, cutting hallucination rates by 30-40%.
- Swapping embedding models without re-indexing silently degrades search quality because embedding spaces are incompatible — version your models explicitly and maintain separate indexes during transitions.
When Retrieval Returns Nothing
Your RAG system works perfectly in testing. You feed it documents, run queries, get relevant chunks back. Deploy to production and suddenly 40% of user queries return empty results — not bad results, literally nothing. The retriever finds zero documents. The LLM falls back to “I don’t have enough information to answer that.”
This doesn’t happen in tutorials because they use clean, preprocessed datasets. Production data arrives messy, inconsistent, and structurally unpredictable.
The problem is embedding normalization drift. During development, you probably embedded your document corpus with consistent preprocessing — lowercase, whitespace normalized, maybe some punctuation stripping. But user queries arrive raw. When query preprocessing doesn’t match document preprocessing, cosine similarity tanks.
Here’s what actually breaks:
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer('all-MiniLM-L6-v2')
# Document embedded during ingestion (preprocessed)
doc_text = "the quick brown fox jumps over the lazy dog"
doc_embedding = model.encode(doc_text, normalize_embeddings=True)
# User query arrives with different capitalization
query_text = "The Quick Brown Fox Jumps Over The Lazy Dog"
query_embedding = model.encode(query_text, normalize_embeddings=True)
# Cosine similarity
similarity = np.dot(doc_embedding, query_embedding)
print(f"Similarity: {similarity:.4f}") # Expected ~1.0, often gets 0.92-0.95
# Now with mixed casing and punctuation
query_messy = "Quick brown fox? Jumps!"
query_messy_embedding = model.encode(query_messy, normalize_embeddings=True)
similarity_messy = np.dot(doc_embedding, query_messy_embedding)
print(f"Messy similarity: {similarity_messy:.4f}") # Drops to 0.85-0.88
The embeddings are close but not identical. When you’re ranking hundreds of documents and your similarity threshold is 0.9, these queries suddenly return nothing.
The fix isn’t just “normalize everything.” You need deterministic preprocessing in both directions:
import re
from functools import lru_cache
class ConsistentPreprocessor:
def __init__(self):
# Cache compiled regex patterns
self.whitespace_pattern = re.compile(r'\s+')
self.punctuation_pattern = re.compile(r'[^\w\s]')
def preprocess(self, text: str) -> str:
# Order matters — test both directions
text = text.lower()
text = self.punctuation_pattern.sub(' ', text)
text = self.whitespace_pattern.sub(' ', text)
text = text.strip()
# This catches the edge case that breaks everything:
# Unicode normalization (NFC vs NFD)
import unicodedata
text = unicodedata.normalize('NFC', text)
return text
@lru_cache(maxsize=10000)
def preprocess_cached(self, text: str) -> str:
# Cache hit rate in production: ~60% for repeated queries
return self.preprocess(text)
preprocessor = ConsistentPreprocessor()
# Now both sides match
doc_clean = preprocessor.preprocess(doc_text)
query_clean = preprocessor.preprocess(query_text)
print(f"Doc: '{doc_clean}'")
print(f"Query: '{query_clean}'")
print(f"Match: {doc_clean == query_clean}") # True
But here’s the part no tutorial mentions: you need to version your preprocessing logic alongside your embeddings. When you refine preprocessing (and you will — maybe you discover that stemming helps, or that keeping certain punctuation is critical for code snippets), you have to re-embed your entire corpus or maintain parallel indexes.
We learned this the hard way after shipping a “minor preprocessing improvement” that silently broke retrieval for 20% of our document base. The symptoms looked like database corruption because the same queries that worked yesterday returned nothing today.

Chunk Boundaries Destroy Context
Every RAG tutorial tells you to chunk documents. The magic numbers vary — 512 tokens, 1000 tokens, “whatever fits your context window” — but the advice is always the same: split long documents into overlapping chunks, embed each chunk, retrieve the most relevant ones.
Nobody tells you what happens when the answer spans a chunk boundary.
Consider a technical document explaining a configuration file:
document = """
The database connection pool is configured in config.yaml:
pool_size: 20
max_overflow: 10
timeout: 30
These settings control connection reuse. The pool_size parameter
defines the maximum number of connections maintained in the pool.
The max_overflow allows temporary connections beyond pool_size
when all pool connections are in use.
"""
# Naive chunking at 100 characters
chunk_size = 100
chunks = [document[i:i+chunk_size] for i in range(0, len(document), chunk_size)]
for idx, chunk in enumerate(chunks):
print(f"Chunk {idx}: {chunk[:80]}...")
print(f"Length: {len(chunk)}\n")
Output:
Chunk 0: The database connection pool is configured in config.yaml:
pool_size: 20
max_o...
Length: 100
Chunk 1: verflow: 10
timeout: 30
These settings control connection reuse. The pool_size parame...
Length: 100
Chunk 2: ter
defines the maximum number of connections maintained in the pool.
The max_overflow...
Length: 100
User asks: “What does max_overflow do?”
Chunk 0 contains the parameter name but not the explanation. Chunk 1 starts mid-word (“overflowverflow”) and has partial context. Chunk 2 contains the full explanation but doesn’t include the parameter name clearly.
The retriever might return chunk 1 (it contains “max_overflow” in the query) or chunk 2 (it has the explanation), but neither alone gives a complete answer. The LLM generates a hallucinated response or says “I don’t have information about max_overflow” despite the document containing exactly what the user needs.
Sentence-based chunking helps but doesn’t solve it:
import nltk
nltk.download('punkt', quiet=True)
def chunk_by_sentences(text: str, max_sentences: int = 3) -> list:
sentences = nltk.sent_tokenize(text)
chunks = []
current_chunk = []
for sent in sentences:
current_chunk.append(sent)
if len(current_chunk) >= max_sentences:
chunks.append(' '.join(current_chunk))
# Overlap: keep last sentence in next chunk
current_chunk = [current_chunk[-1]]
if current_chunk:
chunks.append(' '.join(current_chunk))
return chunks
sentence_chunks = chunk_by_sentences(document, max_sentences=2)
for idx, chunk in enumerate(sentence_chunks):
print(f"Sentence chunk {idx}:\n{chunk}\n")
This is better — each chunk contains complete sentences — but you still hit cases where the question and answer are in different chunks because they appear in different paragraphs.
The real solution involves contextual chunk boundaries with overlap that preserves semantic units:
from typing import List, Tuple
def smart_chunk(text: str, target_size: int = 500, overlap: int = 100) -> List[Tuple[str, dict]]:
"""
Chunk with semantic boundaries and metadata.
Returns (chunk_text, metadata) tuples.
"""
sentences = nltk.sent_tokenize(text)
chunks = []
current_chunk = []
current_length = 0
start_idx = 0
for idx, sent in enumerate(sentences):
sent_len = len(sent)
if current_length + sent_len > target_size and current_chunk:
# Store chunk with boundary metadata
chunk_text = ' '.join(current_chunk)
metadata = {
'start_sentence': start_idx,
'end_sentence': idx - 1,
'has_next': idx < len(sentences),
'has_prev': start_idx > 0
}
chunks.append((chunk_text, metadata))
# Overlap: include last N characters
overlap_sentences = []
overlap_length = 0
for s in reversed(current_chunk):
if overlap_length + len(s) <= overlap:
overlap_sentences.insert(0, s)
overlap_length += len(s)
else:
break
current_chunk = overlap_sentences + [sent]
current_length = overlap_length + sent_len
start_idx = idx - len(overlap_sentences)
else:
current_chunk.append(sent)
current_length += sent_len
if current_chunk:
chunk_text = ' '.join(current_chunk)
chunks.append((chunk_text, {'start_sentence': start_idx, 'end_sentence': len(sentences)-1, 'has_next': False, 'has_prev': start_idx > 0}))
return chunks
for idx, (chunk, meta) in enumerate(smart_chunk(document, target_size=150, overlap=50)):
print(f"Chunk {idx} (sentences {meta['start_sentence']}-{meta['end_sentence']}):")
print(chunk)
print(f"Boundaries: prev={meta['has_prev']}, next={meta['has_next']}\n")
Now when you retrieve a chunk, you can check the metadata and automatically fetch adjacent chunks if the context seems incomplete. This roughly doubles retrieval latency but cuts hallucination rate by 30-40% in systems where answers frequently span multiple paragraphs.
I’m not entirely sure if there’s a better heuristic for overlap size — 50-100 characters works for prose, but code snippets and structured data (tables, lists) probably need different strategies. The docs claim “semantic chunking” libraries handle this, but in practice we still hit edge cases with nested structures.
Embedding Model Versioning Breaks Indexes Silently
You deploy your RAG system with sentence-transformers/all-MiniLM-L6-v2. Six months later, HuggingFace releases a better model — maybe all-MiniLM-L12-v3 with 5% better retrieval accuracy on your domain. You swap the model in your query pipeline, test a few queries, everything looks good.
Two weeks later, users start complaining that search quality has degraded. Not catastrophically — queries still return results — but relevance is off. Documents that should rank first appear third or fourth. Some queries that worked perfectly now miss obvious results.
The problem: your document embeddings were generated with the old model. Your query embeddings use the new model. The two embedding spaces are not compatible.
This is mathematically obvious once you think about it, but it’s shockingly easy to miss in production:
from sentence_transformers import SentenceTransformer
import numpy as np
# Original model (deployed 6 months ago)
model_v1 = SentenceTransformer('all-MiniLM-L6-v2')
# New model (deployed last week)
model_v2 = SentenceTransformer('all-mpnet-base-v2') # Different architecture
doc_text = "Python decorators modify function behavior without changing source code"
query_text = "how do decorators work in python"
# Original deployment: docs embedded with v1
doc_embedding_v1 = model_v1.encode(doc_text, normalize_embeddings=True)
# After model swap: queries use v2
query_embedding_v2 = model_v2.encode(query_text, normalize_embeddings=True)
# Dimension mismatch if architectures differ
print(f"Doc embedding shape (v1): {doc_embedding_v1.shape}")
print(f"Query embedding shape (v2): {query_embedding_v2.shape}")
# Even if dimensions match, similarity is meaningless
if doc_embedding_v1.shape == query_embedding_v2.shape:
# This computes but the value is garbage
invalid_similarity = np.dot(doc_embedding_v1, query_embedding_v2)
print(f"Invalid cross-model similarity: {invalid_similarity:.4f}")
# Correct comparison: same model for both
query_embedding_v1 = model_v1.encode(query_text, normalize_embeddings=True)
valid_similarity = np.dot(doc_embedding_v1, query_embedding_v1)
print(f"Valid similarity (v1/v1): {valid_similarity:.4f}")
Even when dimensions match (both models output 384-dim vectors), the embedding spaces are trained differently. Cosine similarity between embeddings from different models is mathematically defined but semantically meaningless — like comparing temperatures in Celsius and Fahrenheit by just subtracting the numbers.
The fix requires infrastructure most tutorials skip entirely:
import hashlib
import json
from pathlib import Path
from dataclasses import dataclass
from typing import Optional
@dataclass
class EmbeddingModelVersion:
model_name: str
model_revision: Optional[str]
tokenizer_config: dict
def hash(self) -> str:
# Deterministic hash of model configuration
config_str = json.dumps({
'name': self.model_name,
'revision': self.model_revision,
'tokenizer': sorted(self.tokenizer_config.items())
}, sort_keys=True)
return hashlib.sha256(config_str.encode()).hexdigest()[:8]
class VersionedEmbeddingIndex:
def __init__(self, base_path: Path):
self.base_path = base_path
self.base_path.mkdir(exist_ok=True)
def get_index_path(self, model_version: EmbeddingModelVersion) -> Path:
# Separate index per model version
return self.base_path / f"index_{model_version.hash()}.faiss"
def embed_documents(self, documents: list, model_version: EmbeddingModelVersion):
index_path = self.get_index_path(model_version)
if index_path.exists():
print(f"Using existing index for model {model_version.model_name} ({model_version.hash()})")
# Load from disk
return
print(f"Building new index for model {model_version.model_name} ({model_version.hash()})")
# This is where you'd actually build the FAISS/Pinecone/Chroma index
# Save to index_path
# IMPORTANT: Store model_version hash in index metadata
def query(self, query_text: str, model_version: EmbeddingModelVersion, top_k: int = 5):
index_path = self.get_index_path(model_version)
if not index_path.exists():
raise ValueError(
f"No index found for model {model_version.model_name}. "
f"You need to re-embed documents with this model version."
)
# Load index and perform search
# This catches the mistake at query time instead of silently returning bad results
pass
# Usage
model_v1_config = EmbeddingModelVersion(
model_name='all-MiniLM-L6-v2',
model_revision='main',
tokenizer_config={'max_seq_length': 256}
)
model_v2_config = EmbeddingModelVersion(
model_name='all-mpnet-base-v2',
model_revision='main',
tokenizer_config={'max_seq_length': 384}
)
index = VersionedEmbeddingIndex(Path('./indexes'))
# This forces explicit re-indexing when you change models
index.embed_documents(documents=['doc1', 'doc2'], model_version=model_v2_config)
index.query("test query", model_version=model_v2_config)
This approach requires maintaining multiple indexes during model transitions. For a 10M document corpus with 384-dim embeddings, that’s roughly 15GB of storage per index version. Not trivial, but cheaper than silently degraded search quality.
And here’s the really frustrating part: you can’t catch this with unit tests. Your test suite embeds documents and runs queries using the same model instance, so everything passes. The bug only appears when production query infrastructure uses a different model than the batch indexing pipeline — which might be weeks or months apart.

The Latency Tax Nobody Warns You About
RAG tutorials focus on accuracy metrics — NDCG@10, recall@5, stuff that matters for offline eval. Nobody talks about what happens when retrieval takes 800ms and your SLA is 1 second end-to-end.
The latency breakdown for a typical RAG query:
- Embed query: 20-50ms (CPU) or 5-10ms (GPU)
- Vector search: 100-300ms (depends on index size, architecture)
- Fetch documents: 50-150ms (database/object storage round trip)
- LLM inference: 500-2000ms (depends on model size, input tokens)
Total: 670-2500ms minimum. If your LLM is GPT-4 or Claude via API, add network latency and potential queueing.
I mentioned LLM Memory Calculator: Online Estimators Miss 40% Usage in a previous post — memory pressure directly impacts latency through swapping and cache misses. But even with perfectly tuned memory, retrieval itself is often the bottleneck.
The only reliable fix is aggressive caching with cache warming:
import time
from functools import lru_cache
from collections import OrderedDict
import threading
class RetrievalCache:
def __init__(self, max_size: int = 1000, ttl_seconds: int = 3600):
self.cache = OrderedDict()
self.max_size = max_size
self.ttl_seconds = ttl_seconds
self.lock = threading.Lock()
self.hits = 0
self.misses = 0
def _evict_expired(self):
# This shouldn't happen often but catches stale entries
current_time = time.time()
expired = [k for k, (_, timestamp) in self.cache.items()
if current_time - timestamp > self.ttl_seconds]
for k in expired:
del self.cache[k]
def get(self, query: str):
with self.lock:
if query in self.cache:
result, timestamp = self.cache[query]
if time.time() - timestamp < self.ttl_seconds:
# Move to end (LRU)
self.cache.move_to_end(query)
self.hits += 1
return result
else:
del self.cache[query]
self.misses += 1
return None
def set(self, query: str, result):
with self.lock:
if len(self.cache) >= self.max_size:
# Evict oldest
self.cache.popitem(last=False)
self.cache[query] = (result, time.time())
def hit_rate(self) -> float:
total = self.hits + self.misses
return self.hits / total if total > 0 else 0.0
cache = RetrievalCache(max_size=5000, ttl_seconds=1800)
def retrieve_with_cache(query: str):
# Check cache first
cached = cache.get(query)
if cached is not None:
return cached
# Actual retrieval (expensive)
start = time.time()
# result = your_vector_search(query) # 100-300ms
result = ["doc1", "doc2", "doc3"] # Placeholder
elapsed = time.time() - start
cache.set(query, result)
return result
# Simulate queries
for _ in range(100):
retrieve_with_cache("common query")
retrieve_with_cache("another common query")
retrieve_with_cache("unique query " + str(_))
print(f"Cache hit rate: {cache.hit_rate():.2%}")
print(f"Hits: {cache.hits}, Misses: {cache.misses}")
In production, cache hit rates around 40-60% are typical if you’re serving consumer queries (lots of repetition). For enterprise internal tools, it drops to 10-20% because users phrase questions differently.
But here’s the catch: caching only helps if users repeat queries exactly. “What is max_overflow?” and “what does max_overflow do?” are cache misses even though they’re semantically identical. You need query normalization (which itself adds latency) or semantic caching (embed the query, find similar cached queries, reuse results if similarity > threshold).
Semantic caching adds its own complexity:
import numpy as np
from sentence_transformers import SentenceTransformer
class SemanticCache:
def __init__(self, similarity_threshold: float = 0.95):
self.model = SentenceTransformer('all-MiniLM-L6-v2')
self.cache = [] # List of (query_embedding, result, query_text)
self.threshold = similarity_threshold
def get(self, query: str):
if not self.cache:
return None
query_emb = self.model.encode(query, normalize_embeddings=True)
for cached_emb, result, cached_query in self.cache:
sim = np.dot(query_emb, cached_emb)
if sim >= self.threshold:
# print(f"Cache hit: '{query}' ~= '{cached_query}' (sim={sim:.3f})")
return result
return None
def set(self, query: str, result):
query_emb = self.model.encode(query, normalize_embeddings=True)
self.cache.append((query_emb, result, query))
# Limit cache size (this is naive — production needs better eviction)
if len(self.cache) > 1000:
self.cache.pop(0)
sem_cache = SemanticCache(similarity_threshold=0.92)
# These should hit cache
queries = [
"What does max_overflow do?",
"what is max_overflow",
"explain max_overflow parameter"
]
for q in queries:
cached = sem_cache.get(q)
if cached:
print(f"HIT: {q}")
else:
print(f"MISS: {q}")
sem_cache.set(q, ["result_for_" + q])
Semantic caching cuts cache misses by 20-30% but adds 20-50ms per query for embedding + similarity search. Whether it’s worth it depends on your retrieval latency — if vector search takes 300ms, spending 30ms to avoid it is a good trade. If you’ve already optimized retrieval to 50ms, semantic caching might make things slower.
What about prewarming cache for common queries? Works great if you know what users will ask. Spoiler: you don’t. Even with query logs, the long tail is huge. Our top 100 queries account for maybe 30% of traffic. The remaining 70% is one-off or rare queries that don’t benefit from caching at all. If you’re debugging this at 2am and need a boost, maybe grab some Dark Chocolate Espresso Beans — they’re genuinely helpful when you’re staring at latency percentile graphs trying to figure out why P95 spiked.
Debugging Production RAG
When things go wrong, you need observability. Logging every embedding and search result is too expensive (both storage and latency), but logging nothing means you’re blind.
Minimal viable RAG instrumentation:
import logging
import time
from dataclasses import dataclass, asdict
import json
@dataclass
class RAGQueryMetrics:
query_text: str
query_embedding_time_ms: float
retrieval_time_ms: float
num_results: int
top_result_score: float
llm_time_ms: float
total_time_ms: float
cache_hit: bool
def to_json(self) -> str:
return json.dumps(asdict(self))
logger = logging.getLogger('rag_metrics')
logger.setLevel(logging.INFO)
def query_rag_with_metrics(query: str):
start_total = time.time()
# Embedding
start = time.time()
# query_emb = model.encode(query)
embed_time = (time.time() - start) * 1000
# Retrieval
cache_hit = False
cached = cache.get(query)
if cached:
results = cached
retrieval_time = 0
cache_hit = True
else:
start = time.time()
# results = vector_search(query_emb)
results = [("doc1", 0.89), ("doc2", 0.85)] # Placeholder
retrieval_time = (time.time() - start) * 1000
# LLM
start = time.time()
# response = llm.generate(query, context=results)
llm_time = (time.time() - start) * 1000
total_time = (time.time() - start_total) * 1000
metrics = RAGQueryMetrics(
query_text=query[:100], # Truncate for privacy
query_embedding_time_ms=embed_time,
retrieval_time_ms=retrieval_time,
num_results=len(results),
top_result_score=results[0][1] if results else 0.0,
llm_time_ms=llm_time,
total_time_ms=total_time,
cache_hit=cache_hit
)
logger.info(metrics.to_json())
return "response"
# Usage
query_rag_with_metrics("test query")
Ship these metrics to your monitoring system (Datadog, Grafana, whatever). Alert on:
- P95 retrieval latency > 500ms
- Top result score < 0.7 for >10% of queries
- Cache hit rate < 30%
- Zero results returned for >5% of queries
These thresholds are domain-specific, but the point is to catch degradation before users complain.
When RAG Isn’t the Answer
Sometimes retrieval is the wrong tool. If your queries are mostly “What is X?”-style definitional questions and your document corpus is structured (API docs, glossaries), a simple keyword index might outperform vector search for 80% of queries at 1/10th the latency.
If you’re hitting all three of these problems simultaneously — poor retrieval, context spanning issues, and latency bloat — consider whether fine-tuning a smaller model on your specific domain might be cheaper and faster than maintaining a RAG pipeline. Not every problem needs retrieval augmentation.
FAQ
Q: Can I use multiple embedding models simultaneously to avoid re-indexing?
No — you’d need to maintain separate indexes per model anyway, which is exactly what the versioning approach does. Trying to “blend” embeddings from different models (averaging, concatenating) produces worse results than just using one model consistently.
Q: What’s a good semantic caching similarity threshold?
0.90-0.95 for most use cases. Lower than 0.90 and you’ll get false cache hits where the query intent is actually different. Higher than 0.95 and cache hit rate drops too much to be useful. Test on your actual query distribution — it varies significantly by domain.
Q: Should I use approximate nearest neighbor (ANN) or exact search for RAG retrieval?
ANN (HNSW, IVF) is necessary above ~100k documents. Below that, exact search with FAISS flat index is often faster due to lower overhead. The accuracy loss from ANN is usually negligible (<1% recall difference) if you tune the parameters correctly.
What Actually Works
Use consistent preprocessing with versioning. Build overlap into your chunking strategy and include boundary metadata so you can fetch adjacent chunks when needed. Version your embedding models explicitly and maintain separate indexes during transitions. Cache aggressively but measure hit rates — semantic caching helps some workloads and hurts others.
The systems that work in production are the ones that expect failure modes. Retrieval returns nothing sometimes — have a fallback. Chunks miss context — fetch neighbors. Models change — track versions. Latency spikes — cache and monitor.
What I’m still unsure about: whether there’s a better way to handle the semantic caching vs latency tradeoff. In theory, approximate semantic caching (cluster cached queries, map new queries to clusters) should be faster than per-query similarity checks, but I haven’t found a library that does this well. If you’ve solved this, I’d be curious to see the approach.
Did you find this helpful?
Your support keeps this blog running and ad-free content coming.
☕ Buy me a coffeeMost Popular Posts
- Custom Metaclass in Python: 43% Faster Validation (12,838 views)
- Python match-case: 7 Patterns That Beat if-elif Chains (956 views)
- YOLOv8 INT8 Quantization: 4x Faster on Jetson Orin (790 views)
- yfinance Alternatives 2026: 7 Free APIs Compared (753 views)
- PaddleOCR vs EasyOCR vs Tesseract: Why PaddleOCR Is Slower (577 views)