Few-Shot Contamination
Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.
Problem Description
Root Cause: Human operators processing documents don't limit themselves to visible content. They:
- Check other documents in same claim for context
- Look up healthcare procedure codes and standard prices online
- Apply domain knowledge not present in document text
- Make inferences from external systems or databases
Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.
Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."
Production Impact
Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.
Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.
Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.
Solution: Architectural Separation
Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.
Post-Processing Phase: Separate step for enrichment using:
- Business logic and rules
- Cross-document context lookup
- External API calls and database queries
- Domain knowledge application
Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.
Implementation at Scale
alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.
See also
- document-processing-pipeline
- Reference Dataset Design
- LLM Hallucination
- Production AI Systems
- alan-health