Pure Parsing
Document extraction approach that limits LLM output to information directly visible in the source document, excluding external knowledge, cross-document context, or inferred information. Critical architectural pattern developed at alan-health to prevent few-shot-contamination and maintain extraction reliability in production systems.
Core Principle
Document-Grounded Extraction
Constraint: Extract only information explicitly present in the source document
- Visible text: Information that can be read directly from document content
- Visual elements: Data represented in images, tables, stamps, signatures
- No inference: Avoid filling missing information from domain knowledge
- No context: Exclude information from other documents or external sources
Separation from Enrichment
Two-phase architecture: Distinct separation between extraction and business logic
- Pure parsing phase: Document → structured data (document-grounded only)
- Post-processing phase: Structured data + business rules → enriched output
Problem: External Knowledge Contamination
Human Operator Behavior
Manual document processors naturally apply external knowledge:
- Cross-document lookups: Checking related documents in same claim
- Domain expertise: Applying healthcare industry knowledge
- External validation: Confirming procedure codes against standard references
- Business rules: Filling missing fields based on organizational policies
Training Data Corruption
When human-processed examples become few-shot training data:
- Contaminated examples: Training data includes non-document information
- Hallucination learning: LLM learns to invent plausible missing data
- Extraction drift: Model output gradually includes more external knowledge
- Audit trail loss: Cannot verify extracted data against source documents
Solution Architecture
Pure Parsing Implementation
Extraction constraints enforced through:
- Prompt engineering: Explicit instructions to extract only visible information
- Example curation: Few-shot datasets verified to contain only document-grounded data
- Validation rules: Schema checks ensuring extracted fields map to document content
- Human training: Manual processors educated on pure parsing principles
Post-Processing Enrichment
Business logic applied separately:
- Cross-document context: Integration with related documents after pure extraction
- External lookups: API calls to validation services and reference databases
- Domain knowledge: Application of business rules and healthcare expertise
- Inference logic: Filling missing fields based on organizational requirements
Production Benefits
Extraction Reliability
- Hallucination prevention: Eliminates fabricated data not grounded in documents
- Audit transparency: Clear mapping between extracted data and source content
- Quality consistency: Predictable extraction behavior independent of operator knowledge
- Training stability: Few-shot examples remain document-grounded over time
System Modularity
- Independent evolution: Extraction logic can improve without affecting business rules
- Business rule flexibility: Post-processing can change without retraining extraction
- Clear responsibility: Distinct ownership of document processing vs business enrichment
- Testing isolation: Extraction accuracy can be measured independently
Implementation Challenges
Prompt Engineering
Clarity requirements: Instructions must clearly distinguish visible vs inferred information
- Negative examples: Show what NOT to extract from external knowledge
- Boundary cases: Handle situations where document content is ambiguous
- Validation rules: Define exactly what constitutes "visible" information
Example Quality Control
Reference dataset curation: Systematic removal of contaminated training examples
- Human review: Verify examples contain only document-visible information
- Audit trails: Maintain source tracking for all training examples
- Continuous cleaning: Regular review and updating of reference datasets
Performance Impact
Information loss: Some useful enrichment must wait for post-processing phase
- Field completeness: Pure extraction may leave more fields empty
- Processing complexity: Two-phase architecture adds system complexity
- Performance tradeoffs: Multiple processing steps vs single enriched extraction
Validation Techniques
Document Grounding Verification
- Source highlighting: Verify extracted fields can be highlighted in original document
- Human validation: Manual spot-checks of extraction vs document content
- Automated checks: Schema validation ensuring extracted data types match document structure
- Comparative testing: Pure parsing accuracy vs contaminated extraction