~/wiki

Few-Shot Contamination

Mis à jour le 2025-12-22Confiance : high
few-shot-learningllm-hallucinationexample-qualitydocument-processingpure-parsingpost-processing-separationproduction-systemshuman-enriched-examplesexternal-contextknowledge-injectionreference-datasetsmimicry-learningarchitectural-separation

Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.

Problem Description

Root Cause: Human operators processing documents don't limit themselves to visible content. They:

  • Check other documents in same claim for context
  • Look up healthcare procedure codes and standard prices online
  • Apply domain knowledge not present in document text
  • Make inferences from external systems or databases

Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.

Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."

Production Impact

Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.

Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.

Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.

Solution: Architectural Separation

Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.

Post-Processing Phase: Separate step for enrichment using:

  • Business logic and rules
  • Cross-document context lookup
  • External API calls and database queries
  • Domain knowledge application

Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.

Implementation at Scale

alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.

See also