~/wiki

Extrinsic Hallucinations

Mis à jour le 2026-06-11Confiance : high
extrinsic-hallucinationhallucination-typesfactual-accuracyknowledge-groundingllm-reliabilitymodel-limitationsworld-knowledgepre-training-datauncertainty-acknowledgment

A specific type of LLM hallucination where model output is fabricated and not grounded by either the provided context or world knowledge, as opposed to in-context-hallucination which contradicts provided context. It represents a more fundamental challenge in llm-reliability.

Definition and Taxonomy

According to lilian-weng's analysis, extrinsic hallucination occurs when model output should be grounded by the pre-training dataset but is instead fabricated. Since pre-training data serves as a proxy for world knowledge, this essentially means ensuring model output is factual and verifiable by external world knowledge.

More specifically, extrinsic hallucination occurs when language models generate content that is:

  • Fabricated and unfaithful to available knowledge sources
  • Not grounded in the pre-training dataset
  • Not verifiable against external world knowledge
  • Generated with inappropriate confidence despite lack of supporting evidence

This contrasts with in-context-hallucination, where the model output should be consistent with source content explicitly provided in context.

Core Requirements

To avoid extrinsic hallucination, LLMs need to be:

  1. Factual (factual-accuracy): Outputs must align with and be grounded in verifiable world knowledge
  2. Epistemically humble (uncertainty-acknowledgment): Acknowledge uncertainty when knowledge is incomplete

The model should explicitly say "I don't know" when it lacks sufficient information rather than fabricating confident but potentially false responses.

Mitigation and Evaluation Challenges

Detecting and preventing extrinsic hallucination presents unique difficulties:

  • Scale problem: Given the massive size of pre-training datasets, it's computationally expensive to retrieve and identify conflicts per generation. Pre-training datasets are too large to efficiently retrieve and verify against.
  • World knowledge proxy: Pre-training data serves as an approximate representation of world knowledge.
  • Verification complexity: Determining what constitutes "factual" information requires external knowledge sources.

This makes extrinsic hallucination particularly challenging to detect and prevent compared to in-context hallucinations where the source material is explicitly available.

Relationship to Other Concepts

Practical Implications

Understanding extrinsic hallucination is essential for:

  • Developing robust evaluation methodologies
  • Implementing appropriate uncertainty estimation
  • Building trustworthy AI systems for knowledge-intensive tasks
  • Designing training objectives that encourage epistemic humility

See also