Applied AI / 5 MIN READ

Debug a weak RAG answer in order

A practical investigation sequence that separates evidence failures from answer failures.

Keep one failing query fixed while investigating. Record the expected evidence and what the system actually retrieved. If the necessary document is absent from the index, changing the final prompt will not restore it. Check parsing, chunk boundaries, embedding version, metadata filters and document freshness.

Next inspect candidates before and after reranking. A relevant source may be retrieved but dropped by a small candidate budget, an aggressive reranker or a context-packing decision. Compare each stage on the same labeled queries instead of treating an end-to-end score as a diagnosis.

If sufficient evidence reaches the prompt, examine whether the generator uses it faithfully. Separate missing claims, unsupported additions, citation errors and reasonable abstention. A citation ID is useful provenance, but its presence does not prove the cited passage supports the sentence.

Retest the repair on a representative regression set. One successful example demonstrates a local improvement; it does not establish overall quality. Report changes in answer success, access-control failures, latency and cost together. Preserve the failing example as a regression case without exposing private source text to an unrelated audience.