RAG pipelines in production: what breaks after the demo
Every RAG demo looks the same: a folder of PDFs, a vector store, a chat window, and the model answers beautifully. The client is impressed. Then the pipeline meets real data, real traffic and a real budget - and starts falling apart in four predictable places.
I've shipped RAG systems for document-heavy products: internal knowledge bases, support assistants, contract analysis. The details always differ; the failure modes almost never do. This post is the checklist I wish someone had handed me before my first production deployment - and the one I now run through before writing any pipeline code.
1. Chunking: the demo lies to you
In the demo you chunked a few clean documents. In production you get scanned PDFs, tables, code blocks, headers that mean nothing without context. A chunk that made sense in isolation becomes garbage once it's split mid-table or mid-sentence - and the retrieval layer will happily serve that garbage to the model, which will confidently build an answer on top of it.
The fix is unglamorous: chunk by document structure, not by token count.
# structure-aware chunking (pseudocode)
for section in document.sections:
chunks = split(section.text,
max_tokens=500,
keep_together=["table", "code"])
for chunk in chunks:
chunk.metadata = {
"title": document.title,
"path": section.breadcrumb, # "Billing → Refunds → EU"
"updated": document.updated_at
}
What actually helps:
- Chunk by document structure first (headings, sections), not by token count.
- Keep tables and code blocks whole - split around them, not through them.
- Add context to every chunk: document title, section path, date. Retrieval quality jumps more from this than from any embedding model upgrade.
A concrete example from a support-assistant project: we stored the section path with every chunk ("Billing → Refunds → EU customers"), and questions like "can I downgrade my plan?" started hitting the right documents immediately. Same embeddings, same model - just smarter chunk boundaries and metadata. That single change moved our eval scores more than the embedding upgrade we'd been debating for a week.
2. Retrieval quality degrades silently
The demo had 50 documents. Production has 50,000. The same top-5 retrieval that felt magical on a small corpus now returns plausible-looking but wrong chunks - and nobody notices for weeks, because the model still produces confident, well-formatted answers. The system doesn't crash. It just gets slowly, invisibly worse.
This is the most dangerous failure mode precisely because it's invisible. A crashed pipeline gets fixed in an hour. A quietly degrading one erodes trust for a quarter, and by the time someone says "the bot has been useless lately", you have no idea which of the last twenty changes did it.
Fix: build a small eval set early. 30-50 real questions with known answers, run on every pipeline change. It takes an afternoon to build and saves months of "why is it suddenly wrong" debugging.
Two practical notes:
- Take the questions from real users, not from your imagination. The first version of our eval set was full of questions I would ask. Real users asked about billing edge cases and password resets - things I never thought to test.
- Track recall@k (did the right chunk make it into the top-k?), not vibes. It's one number, it's comparable across weeks, and it turns "feels worse" arguments into data.
3. Cost grows non-linearly
Demo: 100 queries a day. Production: the same query asked 40 times a day by different users, plus agents that retry on failure, plus embeddings regenerated on every document update - including the updates nobody asked for.
What keeps the bill sane:
- Cache embeddings. They change only when the document changes. Keyed by content hash, they're the cheapest win in the whole pipeline.
- Cache answers. Even a simple TTL cache on (query, filters) cuts 30-50% of LLM calls in support-style products, where the same questions repeat all day.
- Route easy queries to a smaller model. A cheap classifier (or just the retrieval scores themselves) decides: high-confidence retrieval + short answer → small model; anything ambiguous → the big one.
Numbers from one project: answer caching alone cut LLM spend by ~40% in the first week. The routing rule was crude - retrieval score above a threshold went to the small model - but it worked, and the eval set told us quality hadn't moved. Without evals we would never have dared to turn that knob.
4. No evals = no product
The hardest part of RAG in production is not retrieval or generation - it's knowing whether last week's change made things better or worse. Without evals you're tuning by vibes, and vibes regress.
Minimum viable evals:
- a golden set of questions with known-good answers (from real users, see above);
- an LLM-as-judge for answer quality - with a rubric ("does the answer contain X? is it grounded in the retrieved chunks?"), not a vague "is this good?";
- a nightly run that compares scores against the previous version and posts the diff somewhere the team actually reads.
That's a day of work that makes every future change measurable. The judge doesn't have to be perfect - it has to be consistent. A judge that's wrong the same way every night is still good enough to tell you that Tuesday's chunking change made Wednesday worse.
5. The boring stuff that kills RAG projects
Three things that never appear in demos but decide whether the system survives its first quarter:
- Permissions. If a user shouldn't see a document, retrieval must respect that - filter at query time, inside the vector search, not after generation. Filtering "after the fact" means the model has already seen the document, and prompt-please-forget-it is not a security boundary.
- Stale data. A knowledge base that updates weekly needs a re-index pipeline that handles deletions, not just additions. The worst production bug I've seen in this space was an answer quoting a pricing page from two versions ago - technically correct retrieval, factually wrong product.
- Ownership. Someone must own the pipeline: watch the eval scores, the costs, the index freshness. "The AI does it" is not an owner. Unowned RAG systems decay exactly like unowned databases - just less obviously.
What I'd do differently
If I started a RAG project today, I'd spend the first week not on the pipeline but on the eval set and the chunking strategy. The pipeline itself is a weekend of work; making it reliably good is the actual project. Every hour spent on evals and chunk metadata pays itself back the first time someone asks "did the last change make it better or worse?" - and you can answer with a number instead of a shrug.
Building something with LLMs and want it to survive production? Get in touch - I've shipped RAG pipelines that hold under real load.