~/wiki

Pure Parsing

Confiance : high
pure-parsingdocument-extractionexternal-knowledge-separationhallucination-preventionfew-shot-contaminationpost-processing-separationalan-healthproduction-systemsdocument-grounded-extractionpipeline-architecture

Document extraction approach that limits LLM output to information directly visible in the source document, excluding external knowledge, cross-document context, or inferred information. Critical architectural pattern developed at alan-health to prevent few-shot-contamination and maintain extraction reliability in production systems.

Core Principle

Document-Grounded Extraction

Constraint: Extract only information explicitly present in the source document

  • Visible text: Information that can be read directly from document content
  • Visual elements: Data represented in images, tables, stamps, signatures
  • No inference: Avoid filling missing information from domain knowledge
  • No context: Exclude information from other documents or external sources

Separation from Enrichment

Two-phase architecture: Distinct separation between extraction and business logic

  1. Pure parsing phase: Document → structured data (document-grounded only)
  2. Post-processing phase: Structured data + business rules → enriched output

Problem: External Knowledge Contamination

Human Operator Behavior

Manual document processors naturally apply external knowledge:

  • Cross-document lookups: Checking related documents in same claim
  • Domain expertise: Applying healthcare industry knowledge
  • External validation: Confirming procedure codes against standard references
  • Business rules: Filling missing fields based on organizational policies

Training Data Corruption

When human-processed examples become few-shot training data:

  • Contaminated examples: Training data includes non-document information
  • Hallucination learning: LLM learns to invent plausible missing data
  • Extraction drift: Model output gradually includes more external knowledge
  • Audit trail loss: Cannot verify extracted data against source documents

Solution Architecture

Pure Parsing Implementation

Extraction constraints enforced through:

  • Prompt engineering: Explicit instructions to extract only visible information
  • Example curation: Few-shot datasets verified to contain only document-grounded data
  • Validation rules: Schema checks ensuring extracted fields map to document content
  • Human training: Manual processors educated on pure parsing principles

Post-Processing Enrichment

Business logic applied separately:

  • Cross-document context: Integration with related documents after pure extraction
  • External lookups: API calls to validation services and reference databases
  • Domain knowledge: Application of business rules and healthcare expertise
  • Inference logic: Filling missing fields based on organizational requirements

Production Benefits

Extraction Reliability

  • Hallucination prevention: Eliminates fabricated data not grounded in documents
  • Audit transparency: Clear mapping between extracted data and source content
  • Quality consistency: Predictable extraction behavior independent of operator knowledge
  • Training stability: Few-shot examples remain document-grounded over time

System Modularity

  • Independent evolution: Extraction logic can improve without affecting business rules
  • Business rule flexibility: Post-processing can change without retraining extraction
  • Clear responsibility: Distinct ownership of document processing vs business enrichment
  • Testing isolation: Extraction accuracy can be measured independently

Implementation Challenges

Prompt Engineering

Clarity requirements: Instructions must clearly distinguish visible vs inferred information

  • Negative examples: Show what NOT to extract from external knowledge
  • Boundary cases: Handle situations where document content is ambiguous
  • Validation rules: Define exactly what constitutes "visible" information

Example Quality Control

Reference dataset curation: Systematic removal of contaminated training examples

  • Human review: Verify examples contain only document-visible information
  • Audit trails: Maintain source tracking for all training examples
  • Continuous cleaning: Regular review and updating of reference datasets

Performance Impact

Information loss: Some useful enrichment must wait for post-processing phase

  • Field completeness: Pure extraction may leave more fields empty
  • Processing complexity: Two-phase architecture adds system complexity
  • Performance tradeoffs: Multiple processing steps vs single enriched extraction

Validation Techniques

Document Grounding Verification

  • Source highlighting: Verify extracted fields can be highlighted in original document
  • Human validation: Manual spot-checks of extraction vs document content
  • Automated checks: Schema validation ensuring extracted data types match document structure
  • Comparative testing: Pure parsing accuracy vs contaminated extraction

See also