~/wiki

In-Context Hallucination

Mis à jour le 2025-01-04Confiance : high
in-context-hallucinationhallucination-typescontext-consistencyreading-comprehensionllm-reliabilityfaithfulnesssource-context

A type of LLM hallucination where model output contradicts or is inconsistent with information explicitly provided in the input context, representing a failure in reading comprehension and faithfulness to source material.

Definition

In-context hallucination occurs when language models generate content that:

  • Contradicts explicitly provided context information
  • Shows inconsistency with source content in the input
  • Demonstrates failure in reading comprehension
  • Produces outputs that are unfaithful to the given materials

This is distinct from extrinsic-hallucination, where outputs are fabricated and not grounded in world knowledge or training data.

Core Characteristics

According to lilian-weng's taxonomy:

  • Context dependence: The evaluation baseline is the provided context, not external knowledge
  • Faithfulness requirement: Model outputs should be consistent with source content
  • Reading comprehension failure: Represents a fundamental inability to accurately process given information

Evaluation Approach

In-context hallucination is typically easier to evaluate than extrinsic-hallucination because:

  • The source material is explicitly provided
  • Verification can be done against known context
  • No external knowledge retrieval required
  • Clear ground truth for comparison

Relationship to Other Concepts

Practical Implications

Understanding in-context hallucination is crucial for:

  • Evaluating reading comprehension capabilities
  • Developing faithful summarization systems
  • Building reliable question-answering applications
  • Implementing context-aware generation controls

Detection Methods

Common approaches include:

  • Automated consistency checking against provided context
  • Human evaluation of faithfulness to source material
  • Semantic similarity measurement between output and input
  • Contradiction detection using natural language inference

See also