~/wiki

Classification Bottleneck

Mis à jour le 2025-12-22Confiance : high
classification-bottleneckdocument-processingsingle-point-failureschema-mismatchconcatenated-documentsmisclassification-impacthealthcare-documentsfrench-documentsclassification-accuracyhospital-attestationsemergency-invoicesvocabulary-overlap

Critical single point of failure in document processing pipelines where incorrect document classification leads to wrong extraction schema application, making results unusable. Major production challenge identified at alan-health processing French healthcare documents, where misclassification renders subsequent extraction steps ineffective regardless of extraction quality.

Problem Definition

Single Point of Failure: If classifier predicts wrong document category, extraction step applies incorrect schema to document, producing unusable structured output.

Chain Reaction: Even perfect extraction logic fails when operating on wrong document type assumptions - attempting to extract invoice fields from prescription document structure.

Quality Amplification: Classification errors amplify downstream, turning high-quality extraction capabilities into garbage output due to schema mismatch.

Specific Failure Modes

Vocabulary Overlap: French hospital attestations frequently misclassified as emergency invoices due to shared medical terminology and similar document structure patterns.

Concatenated Documents: Users upload multiple document types in single PDF (prescription + invoice + payment receipt), confusing classifiers designed to expect one document type per upload.

Subtle Distinctions: Similar document layouts with different business purposes require nuanced classification that current systems struggle with at production scale.

Production Impact at Scale

alan-health identifies classification accuracy as limiting factor for overall pipeline performance:

  • Perfect extraction becomes worthless with wrong classification
  • Human review required for misclassified documents regardless of extraction confidence
  • Single bottleneck preventing automation rate improvements

Mitigation Strategies

Multi-Modal Classification: Leverage both text content and visual layout for improved category distinction.

Confidence Thresholding: Route low-confidence classifications to human review before extraction attempt.

Document Splitting: Detect and handle concatenated documents through layout analysis or content segmentation.

Active Learning: Continuously improve classifier on production misclassification examples.

See also