RAG in Production: Stale Indexes, Broken Retrieval, and the Observability You Need


RAG in Production: Stale Indexes, Broken Retrieval, and the Observability You Need

Almost everyone builds their first RAG system the same way. Embed a folder of PDFs, drop the vectors into a local Chroma instance, wire it up with LangChain, ask a question, get a plausible answer. Ship it.

Then it goes to production and degrades in ways that are genuinely hard to see. Nothing throws. Latency looks fine. The answers are fluent, confident, and increasingly wrong.

This post is about the part after the demo: how retrieval actually works underneath, what it takes to keep an index correct as the underlying documents change, and how to build enough observability that “why did it retrieve that?” is a question you can answer in two minutes instead of two days.

Everything below uses one running system as the example. Atlas is the support copilot at a healthcare billing SaaS. Its corpus is about 42,000 documents:

  • ~1,800 help center articles, edited constantly
  • ~600 payer policy PDFs, versioned, occasionally retracted outright
  • ~400 internal runbooks
  • ~39,000 resolved support tickets, append-only

Support agents ask Atlas things like “what’s the refund window on a resubmitted claim?” and it answers from the corpus. That question will come back to haunt us later in this post, because Atlas got it wrong for eleven weeks and nobody noticed.

How the machinery actually works

The premise is that instead of asking a model to answer from memory, you retrieve relevant text at query time and put it in the prompt. The model stops being a knowledge store and becomes a reasoner over supplied context. That’s the whole trick, and it’s why RAG became the default way to ground an LLM in proprietary or fast-moving knowledge.

What matters operationally is that this is two separate systems on two separate clocks.

INDEXING PIPELINE  (offline, batch, minutes to hours)
────────────────────────────────────────────────────────────
 raw docs ──► chunk ──► embed ──► store (vectors + metadata + text)
                                        │
                                        ▼
                                 ┌─────────────┐
                                 │   vector    │
                                 │   index     │
                                 └─────────────┘
                                        ▲
QUERY PIPELINE  (online, per request, milliseconds)          │
────────────────────────────────────────────────────────────│
 question ──► embed ──► search ──┘
                          │
                          └──► rerank ──► assemble prompt ──► LLM ──► answer

Nearly every production incident in RAG comes from these two pipelines disagreeing about the state of the world: the index believes something the corpus no longer does.

The retrieval math, briefly

Retrieval ranks chunks by cosine similarity — the cosine of the angle between the query vector and each chunk vector. Magnitude is ignored, direction is everything:

                 q · d
sim(q, d) = ─────────────────      ranges from -1 to 1
              ‖q‖ · ‖d‖            in practice, 0 to 1 for text embeddings

Computing this against every vector is exact and, past a few million chunks, far too slow. So vector databases use approximate nearest neighbor search. The dominant algorithm is HNSW, which builds a layered proximity graph at index time: sparse long-range links on top for coarse navigation, dense short-range links at the bottom for precision. Search descends the layers greedily and lands in roughly O(log n).

“Approximate” is a real word here, and it has a knob. ef_search controls how many candidates the graph traversal keeps in play:

ef_search = 40    →  ~3ms    recall@10 ≈ 0.88   ← misses roughly 1 in 8 true neighbors
ef_search = 128   →  ~9ms    recall@10 ≈ 0.97
ef_search = 512   →  ~35ms   recall@10 ≈ 0.995

Those numbers move with your data, but the shape holds. The thing to internalize: a chunk can be perfectly indexed and still never be retrieved, because the graph walk didn’t go that way. If you’re debugging a “why wasn’t this found” report, check ef_search before you blame the chunking. Then check the chunking, because it’s usually the chunking.

Chunking: where RAG quietly fails

Chunks need to be small enough to be specific and large enough to contain a complete thought. That sentence is easy to write and hard to satisfy, and getting it wrong produces the most expensive class of bug in RAG — the kind where retrieval succeeds, the answer is confident, and it’s wrong.

Here’s Atlas’s actual failure, which is worth walking through in full because it is extremely typical.

The source document is policy-refunds-4.2.pdf. Naive 512-token fixed-size chunking with 128-token overlap split it like this:

──── chunk 7 ─────────────────────────────────────────────────────
  Refunds for overpaid claims are issued to the originating account
  within 30 days of the original charge. Partial refunds follow the
  same schedule. Refund status is visible in the Billing tab...
───────────────────────────────────────────────────────────────────

──── chunk 8 ─────────────────────────────────────────────────────
  Exceptions:
  (a) claims flagged for payer review are held until adjudication
      completes, which may exceed 90 days;
  (b) accounts in collections are ineligible until the balance
      is settled;
  (c) resubmitted claims restart the window from the resubmission
      date, not the original charge date.
───────────────────────────────────────────────────────────────────

An agent asks: “what’s the refund window on a resubmitted claim?”

Chunk 7 scores 0.87. It’s dense with the query’s vocabulary — “refund”, “window”, “30 days”, “claim”. Chunk 8 scores 0.61, because chunk 8 never says the word “refund.” It says “Exceptions,” and the noun it modifies lives in a different chunk. The retriever does exactly what it was built to do and returns chunk 7. The model reads chunk 7, which is accurate, complete-sounding, and directly on topic, and answers: “30 days from the original charge.”

The correct answer is clause (c): 30 days from the resubmission date. Atlas was off by up to several weeks on a question agents asked dozens of times a week, and there was no error anywhere in the system to find. Faithfulness was perfect. The model didn’t hallucinate a thing. It faithfully reasoned over the wrong context.

The fix wasn’t a better model or more top_k. It was carrying structure into the chunks:

──── chunk 8 (structure-aware) ────────────────────────────────────
  [Refund Policy 4.2 › Section 3: Refund Timing › Exceptions]

  Exceptions to the 30-day refund window:
  (a) claims flagged for payer review are held until adjudication...
  (c) resubmitted claims restart the window from the resubmission
      date, not the original charge date.
───────────────────────────────────────────────────────────────────

Now the chunk contains “refund” and “window” in its own text, its score jumps to 0.84, and it retrieves alongside chunk 7. The model sees both and answers correctly.

The strategies that survive contact with production

Recursive splitting. Split on paragraph breaks first, then sentences, then raw characters only as a last resort. Cheap, no model calls, and dramatically better than counting characters. This should be your floor, not your ceiling.

Semantic chunking. Embed consecutive sentences and cut where similarity between neighbors drops:

sentence pairs, cosine similarity between adjacent embeddings

  s1→s2  0.91  ─┐
  s2→s3  0.88   │  same topic: claim submission timing
  s3→s4  0.90  ─┘
  s4→s5  0.42  ═══  ✂ topic shift — cut here
  s5→s6  0.87  ─┐
  s6→s7  0.89  ─┘   new topic: appeals process

This finds genuine topic boundaries instead of arbitrary positions. It costs one embedding pass over the corpus at index time, which is usually worth it for prose-heavy documents where structure is weak.

Structure-aware splitting. The highest-leverage option when your documents have structure worth respecting:

CorpusSplit atAlways carry along
Legal / policyclause and sub-clause boundariesfull heading path
Codefunction and class boundaries via ASTfile path, imports, class signature
Markdown docsheading hierarchybreadcrumb of parent headings
Ticketsone ticket = one chunk, never splitproduct area, resolution status
Tablesone row group per chunkthe header row, repeated

That last column is the part people skip and it’s most of the value. A chunk that has forgotten which document and section it came from is a chunk that can only be retrieved by accident.

Store metadata on every chunk, without exception: source doc_id, section heading path, page number, document version, valid_from timestamp, content hash, and the name and version of the embedding model that produced the vector. Every one of these earns its keep later — for filtering, for staleness detection, for debugging, and for the migration you haven’t planned yet.

Embedding models and the lock-in you’re signing up for

The embedding model you pick at index time is closer to a schema decision than a library decision. Every vector in your index is an artifact of that specific model. Swap the model and every stored vector becomes geometrically incompatible with your new query vectors — not degraded, incompatible. There is no partial migration. You re-embed everything.

Reasonable production choices as of mid-2026:

ModelDimsNotable property
text-embedding-3-large (OpenAI)3072Best general-purpose recall; supports Matryoshka truncation to 256/1024 dims to trade accuracy for storage
embed-v3 (Cohere)1024Strong multilingual; separate query/document input modes
bge-large-en-v1.5 (BAAI)1024Open weights, self-hostable, competitive on English; no vendor dependency
e5-mistral-7b-instruct4096Instruction-tuned, excellent asymmetric retrieval (short query → long doc); expensive to run

A note on dimensions, since it’s usually treated as a free parameter: 42,000 documents at 10 chunks each is 420,000 vectors. At 3072 dimensions in float32 that’s about 5.2 GB of raw vectors before the HNSW graph, which typically adds another 30–50%. At 1024 dimensions it’s 1.7 GB. On a hot index that lives in RAM, that difference is the entire infrastructure bill.

The indexing pipeline is the real system

This is where tutorials stop and where production problems start.

Your corpus is not static. Help articles get edited, policies get superseded, payer contracts get retracted, tickets get redacted for PII. If your indexing pipeline can’t express update and delete correctly, your RAG system will serve stale and deleted content with total confidence — which is worse than serving nothing, because nothing is at least visible.

Chunk identity, and why updates aren’t updates

A document that splits into 15 chunks produces 15 vectors with 15 independent IDs. When the document changes, there is no row to UPDATE. You have to:

  1. find all chunk IDs belonging to the old version,
  2. delete those vectors,
  3. re-chunk the new content — which may now produce 17 chunks, or 12,
  4. embed and insert the new set,
  5. record the new mapping.

Vector databases don’t give you step 1. They store vectors, not document lineage. So you keep a document registry — an ordinary Postgres table that owns the truth about what’s in the index:

CREATE TABLE doc_chunk_registry (
    doc_id          TEXT NOT NULL,
    chunk_vector_id TEXT NOT NULL,
    content_hash    TEXT NOT NULL,       -- hash of this chunk's text
    doc_hash        TEXT NOT NULL,       -- hash of the whole source doc
    version         INTEGER NOT NULL DEFAULT 1,
    embed_model     TEXT NOT NULL,       -- e.g. 'text-embedding-3-large'
    indexed_at      TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    status          TEXT NOT NULL DEFAULT 'active',  -- active | superseded | deleted
    PRIMARY KEY (doc_id, chunk_vector_id)
);

CREATE INDEX ON doc_chunk_registry (doc_id, status);
CREATE INDEX ON doc_chunk_registry (status, indexed_at);

The registry is the source of truth. The vector store is a derived cache you can rebuild from it. Framing it that way makes the recovery path obvious the first time something goes wrong.

def reindex_document(doc_id: str, new_content: str, vector_store, registry_db):
    # 1. What's currently live for this document?
    old = registry_db.query(
        """SELECT chunk_vector_id FROM doc_chunk_registry
             WHERE doc_id = %s AND status = 'active'""",
        (doc_id,),
    )

    # 2. Remove the old vectors
    vector_store.delete(ids=[r["chunk_vector_id"] for r in old])
    registry_db.execute(
        """UPDATE doc_chunk_registry SET status = 'superseded'
             WHERE doc_id = %s AND status = 'active'""",
        (doc_id,),
    )

    # 3. Re-chunk and re-embed from scratch
    chunks = splitter.split_text(new_content)
    embeddings = embed(chunks)
    new_ids = vector_store.upsert(embeddings, metadata=[...])

    # 4. Record the new mapping
    for cid, chunk in zip(new_ids, chunks):
        registry_db.execute(
            """INSERT INTO doc_chunk_registry
                 (doc_id, chunk_vector_id, content_hash, doc_hash,
                  version, embed_model)
               VALUES (%s, %s, %s, %s, %s, %s)""",
            (doc_id, cid, sha256(chunk), sha256(new_content),
             next_version, EMBED_MODEL),
        )

Two things about this code are load-bearing.

It is not transactional, and it cannot be. The vector store and Postgres are separate systems with no shared commit. If the process dies after step 2 and before step 4, you have vectors in the index with no registry entry — orphans that will be retrieved forever and that nothing will ever clean up, because nothing knows they exist. You need a reconciler:

# nightly: anything in the index the registry doesn't claim is an orphan
def reconcile(vector_store, registry_db):
    indexed = set(vector_store.list_ids())
    claimed = set(registry_db.query_scalar(
        "SELECT chunk_vector_id FROM doc_chunk_registry WHERE status = 'active'"
    ))

    orphans = indexed - claimed          # in index, not in registry → delete
    missing = claimed - indexed          # in registry, not in index → re-embed

    metrics.gauge("rag.index.orphans", len(orphans))
    metrics.gauge("rag.index.missing", len(missing))
    return orphans, missing

Alert on both gauges. A nonzero, growing orphan count is your index rotting in real time.

Delete is the operation that actually bites. When legal retracts payer-policy-aetna-2024.pdf, “delete” has to mean gone — not “we stopped linking to it.” Two traps here. First, most vector engines soft-delete: the vector is tombstoned and excluded from results, but the bytes remain until a compaction or rebuild. Fine for correctness, not fine if your obligation is actual erasure. Know which one your engine does, and know when compaction runs. Second, deletes fail silently far more often than inserts do, because nobody asserts on them. Verify:

def hard_delete_document(doc_id, vector_store, registry_db):
    ids = registry_db.query_scalar(
        "SELECT chunk_vector_id FROM doc_chunk_registry WHERE doc_id = %s", (doc_id,)
    )
    vector_store.delete(ids=ids)

    # Trust nothing. Confirm they're actually unreachable.
    still_there = vector_store.fetch(ids=ids)
    if still_there:
        raise IndexIntegrityError(
            f"{len(still_there)}/{len(ids)} chunks of {doc_id} survived delete"
        )

    registry_db.execute(
        "UPDATE doc_chunk_registry SET status = 'deleted' WHERE doc_id = %s", (doc_id,)
    )

Don’t re-embed what didn’t change

Re-embedding is the dominant cost in this pipeline. Atlas’s 42,000 documents at ~10 chunks each is 420,000 embedding calls for a full rebuild — hours of wall clock and real money, every time.

The first gate is a content hash on the whole document. In practice most “updates” from a CMS are metadata churn — a title tweak, a reviewer field, a timestamp — that leave the text identical:

def should_reindex(doc_id: str, new_content: str, registry_db) -> bool:
    row = registry_db.query_one(
        """SELECT doc_hash, embed_model FROM doc_chunk_registry
             WHERE doc_id = %s AND status = 'active' LIMIT 1""",
        (doc_id,),
    )
    if row is None:
        return True                                  # never seen it
    if row["embed_model"] != EMBED_MODEL:
        return True                                  # model changed under us
    return sha256(new_content) != row["doc_hash"]     # content actually changed

On Atlas’s nightly sync this gate alone dropped re-embedding from ~11,000 documents per night to ~340.

The second gate is chunk-level hashing, which matters for long, mostly-stable documents. A 200-page payer manual that gets a two-paragraph amendment shouldn’t cost 400 embedding calls:

def incremental_reindex(doc_id, new_content, registry_db, vector_store):
    new_chunks = splitter.split_text(new_content)
    new_hashes = {sha256(c): c for c in new_chunks}

    existing = registry_db.query(
        """SELECT chunk_vector_id, content_hash FROM doc_chunk_registry
             WHERE doc_id = %s AND status = 'active'""",
        (doc_id,),
    )
    existing_hashes = {r["content_hash"]: r["chunk_vector_id"] for r in existing}

    unchanged = new_hashes.keys() & existing_hashes.keys()   # keep as-is
    removed   = existing_hashes.keys() - new_hashes.keys()   # delete
    added     = new_hashes.keys() - existing_hashes.keys()   # embed

    vector_store.delete(ids=[existing_hashes[h] for h in removed])
    if added:
        vector_store.upsert(embed([new_hashes[h] for h in added]), metadata=[...])
    # `unchanged` costs nothing at all

One caveat worth knowing before you build this: it only works if your chunker is deterministic and local. Recursive splitting is — insert a paragraph in section 9 and chunks 1–40 hash identically. Semantic chunking often isn’t; a new sentence can shift a boundary and cascade new hashes through the rest of the document. If you’re doing semantic chunking, measure your actual hit rate before assuming this optimization pays.

Partial updates, and the alias pattern that prevents them

The most underappreciated failure mode in RAG is the half-finished rebuild. You start reindexing 10,000 documents, the job dies at 6,000, and now some documents are at version N and some at N+1 — and the retrieval layer has no way to perceive the seam. Queries return a blend of two worldviews. Answers become subtly inconsistent in ways that are almost impossible to reproduce, because reproducing them depends on which chunk won a similarity contest.

The fix is borrowed straight from Elasticsearch operations: build a whole new index, validate it, swap an alias.

rag_index_2026_08_02   ← previous, kept warm for rollback
rag_index_2026_08_09   ← built overnight, validated against benchmark queries
rag_index_current      ← alias; flips atomically between them

Nothing serves from a partial index because the alias only moves once the build is complete and the gate passes. Make the gate a real one — a fixed set of benchmark questions with known-correct source documents, run against the new index before the swap:

def validate_index(index_name, benchmark) -> bool:
    hits = 0
    for q in benchmark:                        # ~200 curated Q → expected doc_id
        results = search(index_name, q.question, top_k=5)
        if q.expected_doc_id in [r.doc_id for r in results]:
            hits += 1
    recall = hits / len(benchmark)
    log.info(f"{index_name} recall@5 = {recall:.3f}")
    return recall >= 0.92                       # refuse the swap below this

Two hundred curated questions is a weekend of work and it is the single highest-return artifact in the whole system. It converts “the index feels worse this week” into a number that blocks a deploy.

When rebuild latency is unacceptable — the index is enormous, or new content must be live within seconds — use incremental upsert with visibility control instead. Stage vectors with a valid_from in the near future and filter on it at query time, which is the same idea as Postgres MVCC:

# Stage: written now, invisible until the cutover moment
vector_store.upsert(
    vectors=new_embeddings,
    metadata=[{
        "doc_id": doc_id,
        "valid_from": (utcnow() + timedelta(minutes=5)).isoformat(),
        "status": "active",
    } for _ in new_embeddings],
)

# Retrieve: only what's live
results = vector_store.query(
    query_vector=q,
    filter={"valid_from": {"$lte": utcnow().isoformat()}, "status": "active"},
)

Check what that metadata filter costs on your engine before committing to it. Some apply the filter after the ANN walk, which means a restrictive filter can silently shrink your top_k — you asked for 10, the graph returned 10, seven were filtered out, and you’re now answering from three chunks.

Upgrading the embedding model

When a better model ships, every vector you have is wrong in one specific sense: it was produced by a different function, so its position in the space has no defined relationship to vectors from the new model. You cannot query with model B against an index built by model A. The results won’t error. They’ll just be noise wearing the costume of relevance.

So a model upgrade is a full re-embed, and it should be run like a database migration:

  1. Build a shadow index with the new model, in parallel, serving nothing.
  2. Mirror a slice of production queries to it and diff the results — overlap in top_k, position changes, and recall against the benchmark set.
  3. Shift traffic gradually with the alias pattern.
  4. Keep the old index warm until you’re sure.

And add the cheap safeguard that catches this class of bug forever: store embed_model on every chunk and assert on it at query time.

QUERY_MODEL = "text-embedding-3-large"

for chunk in results:
    if chunk.metadata["embed_model"] != QUERY_MODEL:
        metrics.increment("rag.model_mismatch", tags=[f"got:{chunk.metadata['embed_model']}"])
        raise ModelMismatchError(
            f"chunk {chunk.id} embedded with {chunk.metadata['embed_model']}, "
            f"queried with {QUERY_MODEL}"
        )

Without that assertion, a half-migrated index degrades quality steadily and reports nothing. With it, you get a loud, specific error the first time it happens.

Retrieval quality: two fixes worth more than a better model

Two techniques account for most of the accuracy gap between a demo and a system people trust. Both are unglamorous.

Hybrid search, because embeddings are bad at identifiers

Dense vectors encode meaning, which is exactly wrong for the queries where meaning isn’t the point. Atlas gets asked “what does error MB-402 mean?” constantly. Semantically, MB-402 is nearly indistinguishable from MB-401 and MB-407 — the embedding captures “a billing error code,” not which one.

query: "MB-402"

  dense (vector) top 3
    0.79  "Common billing error codes and remediation"     ← generic overview
    0.78  "MB-401: payer ID not recognized"                ← wrong code
    0.77  "MB-407: duplicate claim submission"             ← wrong code

  sparse (BM25) top 3
    14.2  "MB-402: subscriber DOB mismatch"                ← exact term hit
     3.1  "Troubleshooting MB-4xx series errors"
     2.4  "Changelog: MB-402 message text updated"

BM25 is a term-frequency ranking function from the 1990s and it destroys embeddings on exact tokens — codes, SKUs, function names, proper nouns, version strings. Run both and fuse the rankings. Reciprocal Rank Fusion is the standard, it needs no score normalization, and it’s about four lines:

def rrf(dense_ranking, sparse_ranking, k=60):
    scores = defaultdict(float)
    for rank, doc in enumerate(dense_ranking, start=1):
        scores[doc.id] += 1 / (k + rank)
    for rank, doc in enumerate(sparse_ranking, start=1):
        scores[doc.id] += 1 / (k + rank)
    return sorted(scores.items(), key=lambda kv: kv[1], reverse=True)

RRF only looks at position, never at the raw scores, which is what makes it safe to combine a cosine similarity of 0.79 with a BM25 score of 14.2 without inventing a normalization scheme you’ll have to maintain.

Reranking, because retrieval and ranking are different jobs

Bi-encoders — what your vector search uses — embed the query and the document separately. That’s what makes them fast enough to search millions of chunks, and it’s also their ceiling: the model never sees the query and the document at the same time, so it can’t reason about how they relate.

A cross-encoder does. It takes (query, chunk) as a single input and scores the pair directly. Vastly more accurate, vastly more expensive — one forward pass per candidate, so it can’t touch the full corpus. Which is exactly why it goes second:

  420,000 chunks
        │  ANN + BM25, fused          ~12ms
        ▼
      top 50
        │  cross-encoder rerank       ~80ms
        ▼
      top 5   ──►  prompt

Retrieve wide and cheap, rank narrow and expensive. On Atlas’s benchmark set, adding a reranker moved recall@5 from 0.83 to 0.94 with no change to chunking, embeddings, or the index. It is the highest ratio of accuracy gained to code written available anywhere in this stack.

Observability: making “why did it say that?” answerable

Production RAG fails in ways that look like model problems and are retrieval problems. The answer is confidently wrong not because the model invented something, but because it reasoned faithfully over the wrong context. From the outside, those two failures are identical. Without tracing, you will spend weeks tuning prompts to fix an indexing bug.

Standard OpenTelemetry traces, metrics and logs are the right substrate, but a RAG request has primitives OTel’s generic span model won’t capture for you. You have to instrument them deliberately.

The span layout

rag_request (root)          trace_id, user_id, question, index_version
  ├── embedding.query       latency, model, input_tokens
  ├── retrieval.dense       latency, top_k, ef_search, num_results
  ├── retrieval.sparse      latency, num_results
  ├── retrieval.fuse        latency, num_candidates
  ├── retrieval.rerank      latency, model, num_in, num_out
  │     └── events: chunk_retrieved × N     ◄── the important part
  ├── prompt.assembly       latency, total_tokens, num_chunks_used, truncated
  └── llm.generate          latency, model, in_tokens, out_tokens, stop_reason

The chunk_retrieved events are what turn a bad answer into a debuggable one. Each event should carry enough to identify the chunk and the version of the world it came from:

for rank, chunk in enumerate(reranked, start=1):
    span.add_event("chunk_retrieved", {
        "rank": rank,
        "chunk_id": chunk.id,
        "doc_id": chunk.metadata["doc_id"],
        "doc_version": chunk.metadata["version"],
        "section": chunk.metadata["heading_path"],
        "dense_score": chunk.dense_score,
        "sparse_score": chunk.sparse_score,
        "rerank_score": chunk.rerank_score,
        "indexed_at": chunk.metadata["indexed_at"],
        "embed_model": chunk.metadata["embed_model"],
        "used_in_prompt": rank <= 5,
    })

Here’s what that bought us. A support agent flags Atlas’s answer to “what’s our refund window?” — it said 60 days, the real answer is 30. Opening the trace:

rag_request  trace=8f2c…  index_version=rag_index_2026_05_31

  retrieval.rerank  →  chunk_retrieved events
    rank 1  doc_id=policy-refunds  version=1  score=0.91  indexed_at=2025-11-02
            section="Refund Policy 1.0 › Timing"
    rank 2  doc_id=policy-refunds  version=1  score=0.88  indexed_at=2025-11-02
    rank 3  doc_id=faq-billing     version=4  score=0.71  indexed_at=2026-05-31

Version 1, indexed nine months ago, when the live document is version 4.2. The answer wasn’t hallucinated — it was a faithful reading of a policy that had been superseded twice. The bug was three layers away, in a reindex job that had been failing its delete step silently since May. Ninety seconds to find, with the trace. Without it, this is a prompt-engineering wild goose chase.

“The system retrieved two chunks from version 1 of a document currently on version 4.2” is an actionable finding. “The system gave a wrong answer” is not.

Logging the why, on a sample

A harder question than “what was retrieved” is “why did the system think that was relevant.” A rerank score of 0.82 doesn’t tell you whether the chunk is genuinely on point or a plausible-looking false positive.

Add a rationale step: after reranking, ask a small model to explain in one line why each of the top 5 chunks is relevant to the query, and attach it to the trace as structured data. This is too expensive to run on every request and too valuable to skip entirely, so sample it:

  • 1% of all production traffic, for baseline drift
  • 100% of user-flagged responses, where you already know something went wrong
  • 100% of requests whose top rerank score falls below a floor — the retriever telling you it wasn’t confident

Closing the loop with an automatic quality signal

The highest-value thing you can build here connects what was retrieved to how good the answer was. That needs an evaluation signal, and for most applications LLM-as-judge is good enough: after generation, send the question, retrieved context, and answer to a small cheap model with a rubric.

Score two things separately, because they fail for different reasons:

  • Faithfulness — did the answer stay inside what the context actually says?
  • Relevance — did the answer address the question that was asked?

Log both against the trace ID. Now you have a queryable dataset, and queries like “every request in the last 7 days where faithfulness < 0.7” return traces you can actually open. Drilling into those, you’ll find they sort into exactly three buckets:

PatternWhat you see in the traceWhere the bug lives
Wrong document entirelydoc_id unrelated to the questionindex corruption, orphans, model mismatch
Right document, wrong sectioncorrect doc_id, wrong heading_pathchunking boundaries
Right context, wrong answercorrect chunks, used_in_prompt=truegeneration — prompt or model

Only chunk-level attribution lets you tell these apart. Without it every bad answer looks the same, and teams default to blaming the model, which is the one explanation that’s usually wrong.

Tag every trace with the index version

One failure mode deserves its own heading because it’s so common and so miserable: the index was updated, retrieval behavior shifted, answer quality dropped, and there’s no way to correlate the two.

span.set_attribute("retrieval.index_version", current_index_alias)
span.set_attribute("retrieval.index_updated_at", index_metadata["updated_at"])
span.set_attribute("retrieval.embed_model", EMBED_MODEL)
span.set_attribute("retrieval.chunker_version", CHUNKER_VERSION)

Four lines. With them, a quality regression becomes a filter — compare mean faithfulness on index_version = new against index_version = old — and the rollback decision takes minutes. chunker_version is there for the same reason: changing your splitter changes retrieval as surely as changing your embedding model does, and it’s even easier to do by accident.

This is obvious in retrospect and almost nobody does it until they’ve spent a bad afternoon trying to explain why quality fell off a cliff on a Tuesday.

Conclusion

Build the benchmark set before the pipeline. Two hundred real questions with known-correct source documents. It gates your index swaps, quantifies every chunking change, and turns “this feels worse” into a number. Every team I’ve seen build this late wished they’d built it first.

Treat the registry as the system of record and the vector index as a cache. Once you can rebuild the index from Postgres in one command, an entire category of incident becomes a chore instead of an emergency.

Assert on deletes. Deletes fail quietly, and a document that was supposed to be gone is the worst thing your system can retrieve — legally as well as operationally. Verify, don’t assume.

Fix chunking before touching models. Every time I’ve been tempted to upgrade the embedding model to fix quality, the actual problem was a chunk that had lost its heading. Structure-aware chunking plus a reranker beats a bigger embedding model, costs less, and doesn’t force a full re-embed.

Put chunk-level attribution in traces on day one. Retrofitting it means retrofitting the ability to debug, and you’ll want it during the first incident, which is precisely when you have no time to build it.

The demo is easy because nothing has changed yet. Production is the state where documents move, policies get retracted, models get deprecated, and somebody asks why the answer was wrong. What separates the two isn’t a better retriever — it’s an indexing pipeline that can express change correctly, a retrieval layer that doesn’t rely on embeddings alone, and enough attribution to tell a retrieval bug from a generation bug in the ninety seconds it should take.