~/wiki

Pure Parsing Separation

Mis à jour le 2025-01-04Confiance : high
pure-parsing-separationdocument-processingarchitectural-designfew-shot-contaminationhallucination-preventionbusiness-logic-separationexternal-context-separationproduction-systemsalan-health

Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.

Problem Addressed

When human operators create training examples, they often include information not visible in the source document:

  • Cross-document context from related claims
  • External knowledge (procedure codes, standard prices)
  • Domain expertise and business logic application

Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.

Architectural Solution

Two-Stage Processing

  1. Pure Parsing Stage: LLM extracts only what's directly visible in the document
  2. Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context

Benefits

  • Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
  • Reduces Hallucinations: LLM cannot invent values based on external context
  • Improves Accuracy: Document extraction becomes more reliable and predictable
  • Enables Validation: Pure parsing results can be verified against source documents

Implementation Requirements

Clean Training Data

  • Review existing few-shot examples for external contamination
  • Separate document-visible information from inferred information
  • Create pure parsing examples that contain only document content

Architectural Separation

  • Independent pure parsing component focused solely on document content
  • Separate post-processing pipeline for enrichment and business logic
  • Clear interface between parsing and enrichment stages

Validation Framework

  • Verify pure parsing outputs against source documents
  • Monitor for hallucination patterns indicating contamination
  • Test enrichment logic independently of parsing accuracy

Production Benefits

alan-health identified this separation as crucial for production reliability:

  • More predictable LLM behavior in document extraction
  • Easier debugging when issues arise (parsing vs enrichment problems)
  • Better evaluation capability for each pipeline stage
  • Reduced maintenance overhead from hallucination issues

See also