Lattice
Skip to content
← Fix a symptom

Band II · State · “wrong answers”

Why does my RAG system hallucinate?

Short answer

RAG hallucinations usually mean the model answered without support in the retrieved text, or the chunks looked relevant but were not. Log retrieval for each failure, fix ranking and chunking before swapping models, and require citations or abstention when context is thin. Build a small eval of questions where every answer must stay inside the passages you retrieved.

By Kunj Shah · Tool facts verified · method

Check in this order

  1. 01

    Score failures where answers must cite retrieved passages

    Layer 09 · Evaluation & Observability

    Without a grader that enforces grounding, you cannot tell whether a prompt change helped or you just got luckier on one demo.

    • promptfoo
      Use when
      Declarative red-teaming and regression tests that run in CI.
      Skip when
      You need a managed UI rather than a test runner.
    • DeepEval
      Use when
      pytest-style evaluation you can run next to your unit tests.
      Skip when
      You need a platform rather than a library.
  2. 02

    Log the retrieved chunks next to each bad answer

    Layer 09 · Evaluation & Observability

    Hallucination in RAG is often 'the model guessed' or 'the wrong chunk ranked first'. You need both sides in one trace to know which.

    • Langfuse
      Use when
      Self-hostable tracing, prompts and evals in one place.
      Skip when
      You want a fully managed product with a support contract.
    • Arize Phoenix
      Use when
      OpenTelemetry-native evaluation you can run yourself.
      Skip when
      You want a vendor support contract behind it.
  3. 03

    Rerank before generation when recall is noisy

    Layer 03 · Retrieval & Vector Stores

    Vector search alone returns plausible neighbours. A cross-encoder second stage is the cheapest way to stop almost-right chunks from steering the model.

    • Rerankers
      Use when
      A cross-encoder in the second retrieval stage, to lift precision cheaply.
      Skip when
      You have no way to measure whether recall actually improved.
  4. 04

    Fix parsing and chunk boundaries on the source documents

    Layer 03 · Retrieval & Vector Stores

    Chunks cut mid-table or mid-sentence produce context that reads well and contains no usable fact — the model fills the gap.

    • Docling
      Use when
      Layout-aware parsing that preserves tables and heading structure.
      Skip when
      Your documents are plain text.
    • Unstructured
      Use when
      Partitioning messy documents into retrievable chunks before indexing.
      Skip when
      Your input is already structured.
  5. 05

    Require citations, abstention, or a structured answer schema

    Layer 08 · Prompt Engineering

    When the fact was in context and still misused, tightening the output contract is cheaper than fine-tuning and easier to revert.

    • Instructor
      Use when
      Schema-validated extraction with typed retry and partial streaming.
      Skip when
      You need a hard token-level guarantee rather than validation.
    • DSPy
      Use when
      Compiling prompts into optimised programs from labelled examples.
      Skip when
      You have no examples and no way to score them.

Looks like a fix, is not

  • Swapping to a larger model before checking what was retrieved for the failures.
  • Stuffing more chunks into the prompt without measuring whether recall improved.
  • Fine-tuning to 'stop hallucinating' when the retrieved passages never contained the answer.

Quick answers

What should I check first?
Score failures where answers must cite retrieved passages. Without a grader that enforces grounding, you cannot tell whether a prompt change helped or you just got luckier on one demo.
What looks like a fix but is not?
Swapping to a larger model before checking what was retrieved for the failures. Stuffing more chunks into the prompt without measuring whether recall improved. Fine-tuning to 'stop hallucinating' when the retrieved passages never contained the answer.