Pure Parsing Separation
Mis à jour le 2025-01-04Confiance : high
pure-parsing-separationdocument-processingarchitectural-designfew-shot-contaminationhallucination-preventionbusiness-logic-separationexternal-context-separationproduction-systemsalan-health
Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.
Problem Addressed
When human operators create training examples, they often include information not visible in the source document:
- Cross-document context from related claims
- External knowledge (procedure codes, standard prices)
- Domain expertise and business logic application
Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.
Architectural Solution
Two-Stage Processing
- Pure Parsing Stage: LLM extracts only what's directly visible in the document
- Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context
Benefits
- Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
- Reduces Hallucinations: LLM cannot invent values based on external context
- Improves Accuracy: Document extraction becomes more reliable and predictable
- Enables Validation: Pure parsing results can be verified against source documents
Implementation Requirements
Clean Training Data
- Review existing few-shot examples for external contamination
- Separate document-visible information from inferred information
- Create pure parsing examples that contain only document content
Architectural Separation
- Independent pure parsing component focused solely on document content
- Separate post-processing pipeline for enrichment and business logic
- Clear interface between parsing and enrichment stages
Validation Framework
- Verify pure parsing outputs against source documents
- Monitor for hallucination patterns indicating contamination
- Test enrichment logic independently of parsing accuracy
Production Benefits
alan-health identified this separation as crucial for production reliability:
- More predictable LLM behavior in document extraction
- Easier debugging when issues arise (parsing vs enrichment problems)
- Better evaluation capability for each pipeline stage
- Reduced maintenance overhead from hallucination issues
See also
- few-shot-contamination
- Document Processing Architecture
- LLM Hallucination Prevention