Band II · State · “wrong answers”
Why does my RAG system hallucinate?
Short answer
RAG hallucinations usually mean the model answered without support in the retrieved text, or the chunks looked relevant but were not. Log retrieval for each failure, fix ranking and chunking before swapping models, and require citations or abstention when context is thin. Build a small eval of questions where every answer must stay inside the passages you retrieved.
By Kunj Shah · Tool facts verified · method
Check in this order
- 01
Score failures where answers must cite retrieved passages
Layer 09 · Evaluation & Observability
Without a grader that enforces grounding, you cannot tell whether a prompt change helped or you just got luckier on one demo.
- 02
Log the retrieved chunks next to each bad answer
Layer 09 · Evaluation & Observability
Hallucination in RAG is often 'the model guessed' or 'the wrong chunk ranked first'. You need both sides in one trace to know which.
- Langfuse
- Use when
- Self-hostable tracing, prompts and evals in one place.
- Skip when
- You want a fully managed product with a support contract.
- Arize Phoenix
- Use when
- OpenTelemetry-native evaluation you can run yourself.
- Skip when
- You want a vendor support contract behind it.
- Langfuse
- 03
Rerank before generation when recall is noisy
Layer 03 · Retrieval & Vector Stores
Vector search alone returns plausible neighbours. A cross-encoder second stage is the cheapest way to stop almost-right chunks from steering the model.
- Rerankers
- Use when
- A cross-encoder in the second retrieval stage, to lift precision cheaply.
- Skip when
- You have no way to measure whether recall actually improved.
- Rerankers
- 04
Fix parsing and chunk boundaries on the source documents
Layer 03 · Retrieval & Vector Stores
Chunks cut mid-table or mid-sentence produce context that reads well and contains no usable fact — the model fills the gap.
- Docling
- Use when
- Layout-aware parsing that preserves tables and heading structure.
- Skip when
- Your documents are plain text.
- Unstructured
- Use when
- Partitioning messy documents into retrievable chunks before indexing.
- Skip when
- Your input is already structured.
- Docling
- 05
Require citations, abstention, or a structured answer schema
Layer 08 · Prompt Engineering
When the fact was in context and still misused, tightening the output contract is cheaper than fine-tuning and easier to revert.
- Instructor
- Use when
- Schema-validated extraction with typed retry and partial streaming.
- Skip when
- You need a hard token-level guarantee rather than validation.
- DSPy
- Use when
- Compiling prompts into optimised programs from labelled examples.
- Skip when
- You have no examples and no way to score them.
- Instructor
Looks like a fix, is not
- Swapping to a larger model before checking what was retrieved for the failures.
- Stuffing more chunks into the prompt without measuring whether recall improved.
- Fine-tuning to 'stop hallucinating' when the retrieved passages never contained the answer.
Quick answers
- What should I check first?
- Score failures where answers must cite retrieved passages. Without a grader that enforces grounding, you cannot tell whether a prompt change helped or you just got luckier on one demo.
- What looks like a fix but is not?
- Swapping to a larger model before checking what was retrieved for the failures. Stuffing more chunks into the prompt without measuring whether recall improved. Fine-tuning to 'stop hallucinating' when the retrieved passages never contained the answer.