~/wiki

Document Processing Pipeline

Confiance : high
document-processingmultimodal-llmocrhealthcare-documentsproduction-systemsevaluation-frameworkclassification-bottleneckpydantic-validationfrench-documentsautomation-ratespipeline-evolutionarticle-series

Production-grade system for extracting structured data from documents using LLMs, particularly effective for complex healthcare and insurance documents. Combines OCR transcription with document images for optimal accuracy, as demonstrated by alan-health's 70% automation rate on French healthcare documents.

Architecture Evolution

Text-Only Processing (Initial Approach)

  • Input: OCR Markdown transcription only
  • Output: Structured data extraction
  • Performance: Good baseline accuracy for most documents

Image-Only Processing (Experimental)

  • Input: Document images only, no OCR
  • Findings: Lower accuracy than text-based extraction
  • Problems: Hallucinations where LLM "read" information not actually on documents
  • Conclusion: OCR text provides more reliable parsing foundation than raw pixel interpretation

Multimodal Processing (Current Best Practice)

  • Input: OCR Markdown transcription + document images combined
  • Performance: Outperforms either input alone
  • Benefits:
    • OCR text provides reliable, precise content parsing
    • Images provide visual layout context, field positioning, table structure
    • Combined approach leverages strengths of both modalities

Production Components

Validation Layer

  • Pydantic schema validation for structured output
  • Failed validation triggers human review with structured error messages
  • Best-effort extraction preserved as starting point for reviewers

Human-in-the-Loop Integration

  • Configurable human review for high-stakes document categories
  • Systematic review for new document categories during rollout
  • Online evaluation comparison with manual parsing results

Reference Dataset Strategy

  • Curated examples bootstrap new document categories
  • Small, hand-picked datasets achieve surprising effectiveness
  • Validated documents become potential few-shot examples over time

Future Evolution

Current multimodal capabilities are maturing rapidly. As models improve image-only processing accuracy, the OCR transcription step may become unnecessary, simplifying the pipeline architecture.

See also