~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Backtest Monitoring

page dédiée →

Production evaluation system that re-runs document processing pipelines on reference datasets with verified ground truth to measure the impact of changes before deployment. Essential safety mechanism implemented at alan-health to prevent silent regressions across document categories.

Core Functionality

Pipeline Re-execution: Complete processing workflow from document input through transcription, classification, and extraction steps on curated reference datasets.

Ground Truth Comparison: Field-by-field comparison between system output and verified expected results across entire extraction schema.

Change Impact Measurement: Quantifies improvements and regressions before changes reach production environment.

Monitoring Components

Classification Diff Analysis:

  • Document category prediction accuracy
  • Sub-class classification performance
  • Category-specific error patterns

Extraction Diff Analysis:

  • Field-level accuracy comparison
  • Value-by-value extraction validation
  • Schema compliance measurement

Criticality Weighting System:

  • High priority: Financial amounts, critical dates, regulatory fields
  • Medium priority: Names, addresses, secondary identifiers
  • Low priority: Formatting variations, optional fields

Dashboard Integration

Visual Monitoring: Real-time tracking of accuracy metrics across document categories with regression alerts and improvement validation.

Aggregate Analysis: Results analyzed per category, per field, or in aggregate to identify patterns and guide development decisions.

Historical Tracking: Longitudinal performance monitoring enabling trend analysis and regression root cause identification.

Production Implementation

Pre-Deployment Validation: Every system change validated against reference datasets before production release.

Automated Safety Net: Prevents deployment of changes that degrade performance on any document category.

Development Confidence: Teams iterate with quantified impact visibility rather than subjective assessment.

Reference Dataset Requirements

Production Authenticity: Real documents from production pipeline with verified manual extractions as ground truth.

Representative Coverage: Spans all document types, quality levels, and edge cases encountered in production environment.

Immutable Standards: Stable reference datasets ensure consistent measurement across evaluation runs.

See also

Clean Architecture

page dédiée →

Software design philosophy emphasizing separation of concerns, dependency inversion, and autonomous components to create maintainable, testable, and evolution-friendly systems. Particularly valuable for production systems requiring long-term maintainability and team handovers.

Core Principles

Separation of Concerns

Each module should have a single, well-defined responsibility:

# Instead of monolithic pipeline.py (1561 lines)
pipeline.py          # Orchestration only (300 lines)
query_processor.py   # Query analysis and reformulation  
retriever.py         # Multi-source data retrieval
section_aggregator.py # Chunk to section transformation
context_builder.py   # Document aggregation logic
generator.py         # Response generation

Dependency Inversion

High-level modules should not depend on low-level modules. Both should depend on abstractions:

# Bad: Direct dependency on specific implementation
from src.rag_v2.embedder import ScalewayEmbedder

# Good: Internal abstraction  
from .embedder import Embedder  # Internalized interface

Autonomous Components

Each major component should be self-contained with minimal external dependencies. This enables independent testing, deployment, and evolution.

Implementation Strategies

Layered Architecture

Presentation Layer: UI components, API endpoints, user interfaces Application Layer: Business logic, orchestration, workflow management
Domain Layer: Core business entities, rules, and domain-specific logic Infrastructure Layer: Database access, external APIs, system interfaces

Component Internalization

Rather than maintaining complex shared libraries, internalize and simplify components for specific use cases:

# Before: General-purpose shared component
class FallbackEmbedder:
    """Supports 5 embedding providers with complex fallback logic"""
    # 425 lines handling multiple providers, retry logic, caching, etc.

# After: Clean, specific implementation  
class Embedder:
    """Albert + Scaleway embeddings for HR document processing"""
    # 200 lines focused on actual requirements

Interface Consistency

Define clear contracts between components that remain stable even as implementations change:

@dataclass
class RetrievalResult:
    chunks: List[DocumentChunk]
    sections: List[DocumentSection]  
    confidence: float
    source_tables: List[str]

class Retriever:
    def retrieve(self, query: str, top_k: int = 20) -> RetrievalResult:
        """Consistent interface regardless of implementation details"""

Real-World Application: RAG System Migration

The assistant-rh clean architecture migration demonstrates these principles in practice:

Before (Unclear Separation):

  • pipeline.py: 1561 lines mixing orchestration, business logic, and data access
  • Complex imports across

Concatenated Document Classification

page dédiée →

Challenge in document processing where users upload multiple document types in a single PDF, confusing classifiers designed to expect one document type per upload. Major production issue identified at alan-health processing French healthcare documents.

Problem Definition

Typical Concatenated Patterns

  • Prescription + invoice + payment receipt in single PDF
  • Hospital attestation + multiple invoices
  • Mixed healthcare document types spanning multiple pages

Classification System Confusion

  • Traditional classifiers trained on single-document-type assumptions
  • Vocabulary overlap between document types creates ambiguity
  • First document in sequence may drive classification decision
  • Extraction schema mismatch leads to unusable results

Impact on Processing Pipeline

Single Point of Failure

  • Incorrect classification renders entire extraction unusable
  • Wrong schema applied to mixed content produces garbled results
  • Human review required even for otherwise processable individual documents

Example: French Healthcare Documents

  • Hospital attestations misclassified as emergency invoices due to shared vocabulary
  • Prescription documents combined with pharmacy invoices confuse categorical boundaries
  • Payment receipts concatenated with medical invoices create classification uncertainty

Potential Solutions

Document Segmentation

  • Pre-processing to identify document boundaries within PDF
  • Individual classification per document segment
  • Parallel extraction pipelines for each identified document type

Multi-Label Classification

  • Update classification system to handle multiple document types
  • Confidence scoring per document type detected
  • Sequential processing based on detected types

Layout-Based Detection

  • Visual layout analysis to identify document boundaries
  • Page-level classification before content-level analysis
  • Structural cues for document type identification

Current Status

alan-health identifies this as active area for improvement in their production system. Classification accuracy improvements needed for these edge cases to reduce human review requirements and improve automation rates.

See also

Document Processing Pipeline

page dédiée →

Production-grade system for extracting structured data from documents using LLMs, particularly effective for complex healthcare and insurance documents. Combines OCR transcription with document images for optimal accuracy, as demonstrated by alan-health's 70% automation rate on French healthcare documents.

Architecture Evolution

Text-Only Processing (Initial Approach)

  • Input: OCR Markdown transcription only
  • Output: Structured data extraction
  • Performance: Good baseline accuracy for most documents

Image-Only Processing (Experimental)

  • Input: Document images only, no OCR
  • Findings: Lower accuracy than text-based extraction
  • Problems: Hallucinations where LLM "read" information not actually on documents
  • Conclusion: OCR text provides more reliable parsing foundation than raw pixel interpretation

Multimodal Processing (Current Best Practice)

  • Input: OCR Markdown transcription + document images combined
  • Performance: Outperforms either input alone
  • Benefits:
    • OCR text provides reliable, precise content parsing
    • Images provide visual layout context, field positioning, table structure
    • Combined approach leverages strengths of both modalities

Production Components

Validation Layer

  • Pydantic schema validation for structured output
  • Failed validation triggers human review with structured error messages
  • Best-effort extraction preserved as starting point for reviewers

Human-in-the-Loop Integration

  • Configurable human review for high-stakes document categories
  • Systematic review for new document categories during rollout
  • Online evaluation comparison with manual parsing results

Reference Dataset Strategy

  • Curated examples bootstrap new document categories
  • Small, hand-picked datasets achieve surprising effectiveness
  • Validated documents become potential few-shot examples over time

Future Evolution

Current multimodal capabilities are maturing rapidly. As models improve image-only processing accuracy, the OCR transcription step may become unnecessary, simplifying the pipeline architecture.

See also

Evaluation Framework Design

page dédiée →

Systematic approach for measuring LLM system changes before production deployment, preventing silent regressions and ensuring quality improvements. Essential for production systems where changes must be validated against real-world performance, as implemented at alan-health with comprehensive backtest monitoring and field-level analysis.

Core Principle: Measure Before You Ship

Pre-Deployment Validation: Every change (new LLM model, updated instructions, different few-shot strategy) measured against reference datasets before reaching production.

Regression Prevention: Avoid scenarios where pharmacy invoice extraction improvements silently degrade hospital invoice accuracy.

Confidence Building: Teams can iterate on extraction instructions, switch models, or adjust strategies with quantified impact visibility.

Reference Dataset Requirements

Ground Truth Quality: Real production documents with verified extractions serving as gold standard for comparison.

Representative Coverage: Documents spanning all categories, edge cases, and quality levels encountered in production.

Versioned and Immutable: Reference datasets remain stable for consistent measurement across evaluation runs.

Evaluation Components

Classification Diff: Measure whether system predicts correct document category and sub-classes compared to expected results.

Extraction Diff: Field-by-field comparison between actual extracted values and expected results across entire schema.

Criticality Weighting: Different importance levels for extraction errors:

  • High criticality: Amount paid, dates, critical financial values
  • Medium criticality: Healthcare professional names, secondary fields
  • Low criticality: Minor formatting or optional field variations

Backtest Execution

Full Pipeline Re-run: Complete processing from document input through transcription, classification, and extraction steps.

Automated Analysis: Systematic comparison storing results for analysis per category, per field, or in aggregate.

Dashboard Monitoring: Visual tracking of accuracy metrics, regression alerts, and improvement validation across document types.

Production Implementation

alan-health uses this framework for:

  • LLM model upgrades and downgrades
  • Prompt engineering iterations
  • Few-shot strategy modifications
  • Pipeline architecture changes
  • Performance optimization validation

Results stored and analyzed to guide development decisions with quantified confidence rather than subjective assessment.

See also

Few-Shot Contamination

page dédiée →

Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.

Problem Description

Root Cause: Human operators processing documents don't limit themselves to visible content. They:

  • Check other documents in same claim for context
  • Look up healthcare procedure codes and standard prices online
  • Apply domain knowledge not present in document text
  • Make inferences from external systems or databases

Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.

Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."

Production Impact

Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.

Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.

Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.

Solution: Architectural Separation

Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.

Post-Processing Phase: Separate step for enrichment using:

  • Business logic and rules
  • Cross-document context lookup
  • External API calls and database queries
  • Domain knowledge application

Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.

Implementation at Scale

alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.

See also

Interface Consistency

page dédiée →

Fundamental software architecture principle requiring that components implementing the same interface return compatible data types and follow identical behavioral contracts. Critical for system reliability and maintainability, particularly in complex systems like RAG pipelines where multiple implementations must be interchangeable.

Core Principle

Interface consistency ensures that:

  • All implementations of an interface return the same data types
  • Method signatures are identical across implementations
  • Behavioral contracts are maintained regardless of underlying implementation
  • Downstream consumers can rely on predictable data formats

Critical Failure Mode: Mixed Return Types

A common violation occurs when different implementations of the same interface return incompatible data types:

# Inconsistent interface implementation
class TFIDFRetriever:
    def search(self, query):
        return [(index, score), (index, score)]  # Returns tuples

class PostgresRetriever:  
    def search(self, query):
        return [Chunk(...), Chunk(...)]  # Returns objects

This creates downstream failures when consumers expect consistent types:

# Downstream consumer crashes with tuples
def build_prompt(chunks):
    for chunk in chunks:
        content = chunk.content  # AttributeError if chunk is tuple

Production Impact

Interface inconsistency can cause:

  • Runtime Crashes: Type errors when unexpected data types are processed
  • Silent Data Corruption: Partial processing with wrong assumptions
  • Debugging Complexity: Failures occur far from the source of inconsistency
  • System Brittleness: Adding new implementations breaks existing code

RAG System Implications

In rag-systems, interface consistency is particularly critical because:

Multiple Retriever Implementations

Systems often support multiple retrieval strategies:

  • Vector-based retrieval returning ranked documents
  • Keyword-based retrieval returning scored matches
  • Hybrid systems combining multiple approaches
  • multi-source-rag federating across different backends

Downstream Processing Chains

Retrieved results flow through multiple processing stages:

  • Reranking systems expecting specific document formats
  • Context builders assembling prompts from document content
  • Logging systems recording search results
  • UI components displaying search results

Pluggable Architecture Requirements

Production RAG systems require:

  • Hot-swappable retriever implementations
  • A/B testing with different retrieval strategies
  • Fallback mechanisms when primary retrievers fail
  • Configuration-driven retriever selection

Design Patterns for Consistency

Common Interface Definition

from abc import ABC, abstractmethod
from typing import List

class Retriever(ABC):
    @abstractmethod
    def search(self, query: str, limit: int = 10) -> List[Chunk]:
        """Return list of Chunk objects ranked by relevance"""
        pass

Adapter Pattern for Legacy Systems

class LegacyRetrieverAdapter(Retriever):
    def __init__(self, legacy_retriever):
        self.legacy = legacy_retriever
    
    def search(self, query: str, limit: int = 10) -> List[Chunk]:
        # Convert legacy tuple format to Chunk objects
        raw_results = self.legacy.search(query, limit)
        return [self._tuple_to_chunk(r) for r in raw_results]

Runtime Type Validation

def validate_search_results(results: List[Chunk]) -> List[Chunk]:
    """Validate that all results are proper Chunk objects"""
    for result in results:
        if not isinstance(result, Chunk):
            raise TypeError(f"Expected Chunk, got {type(result)}")
    return results

Testing Strategies

Interface Compliance Tests

def test_retriever_interface_compliance():
    retrievers = [TFIDFRetriever(), PostgresRetriever(), HybridRetriever()]
    
    for retriever in retrievers:
        results = retriever.search("test query")
        
        # Verify return type consistency
        assert isinstance(results, list)
        for result in results:
            assert isinstance(result, Chunk)
            assert hasattr(result, 'content')
            assert hasattr(result, 'score')

Downstream Integration Tests

def test_build_prompt_with_all_retrievers():
    """Test that build_prompt works with results from any retriever"""
    test_query = "sample query"
    
    for retriever in all_retriever_implementations():
        results = retriever.search(test_query)
        prompt = build_prompt(results)  # Should not crash
        assert isinstance(prompt, str)
        assert len(prompt) > 0

See also

Multimodal Document Processing

page dédiée →

Production approach combining both OCR text transcription and document images as input to LLM systems, outperforming either single-modal approach. Represents evolved best practice as of 2025-2026, demonstrated at scale by alan-health processing French healthcare documents.

Evolution Path

Text-Only → Image-Only → Multimodal

  1. Text-Only Processing: Initial approach using only OCR Markdown transcription as LLM input. Worked well for most documents but lacked visual context.

  2. Image-Only Experiment: Attempted processing with document images alone, skipping transcription. Result: Lower accuracy than text-based extraction, with increased hallucinations where LLMs "read" information not actually present in ambiguous or hard-to-parse images.

  3. Multimodal Combination: Current best practice combining both OCR transcription and document images. Outperforms either input alone.

Why Multimodal Works

Complementary Information Sources:

  • OCR Transcription Provides: Reliable text content that LLMs can parse precisely, reducing ambiguity in character recognition
  • Document Images Provide: Visual layout context, spatial relationships, table structures, presence of stamps/signatures, form organization

Reduced Hallucination Risk: OCR text anchors the extraction to actual document content, while images provide visual verification and context that prevents misinterpretation.

Technical Implementation

Input Format: Send both OCR Markdown transcription AND document image to multimodal LLM models simultaneously.

Processing Strategy: LLM uses text for precise content extraction while referencing image for layout understanding and visual verification.

Production Benefits:

  • Higher extraction accuracy than single-modal approaches
  • Better handling of complex table structures
  • Improved recognition of visual elements (stamps, signatures)
  • Reduced hallucination on ambiguous content

Future Evolution

As of early 2025-2026, multimodal LLMs still require OCR transcription support for optimal accuracy. However, visual capabilities are maturing rapidly - at some point, models may achieve sufficient accuracy with image-only input, potentially eliminating the OCR transcription step.

Production Validation

alan-health's production system processes millions of French healthcare documents using this multimodal approach, achieving 70% automation rates with higher accuracy than previous single-modal systems.

See also

Multimodal LLM Processing

page dédiée →

Advanced document processing approach combining both OCR-extracted text and document images as input to LLMs, achieving superior extraction accuracy compared to either modality alone. Key insight from alan-health: text provides reliable content parsing while images provide crucial visual layout context.

Processing Evolution

Single Modality Limitations

Text-Only Processing: Initial approach using only OCR Markdown transcription worked well for most documents but missed visual context cues.

Image-Only Processing: Experimental approach achieved lower accuracy than text-based extraction, with significant hallucination problems where LLMs would "read" information not actually present in ambiguous or hard-to-parse images.

Optimal Multimodal Approach

Combined OCR + Image: Current best practice providing superior results through:

  • OCR Text: Reliable content that LLMs can parse precisely
  • Document Image: Visual layout context showing field positioning, table structures, stamps, signatures

Technical Advantages

Complementary Information

  • Text provides precise character-level content extraction
  • Images provide spatial relationships and visual formatting context
  • Combined input reduces ambiguity in document structure interpretation

Hallucination Reduction

Image-only processing produced hallucinations where LLMs invented values from ambiguous visual content. OCR text anchors the LLM to actual document content while images provide confirmatory visual context.

Layout Context

Visual elements crucial for accurate extraction:

  • Field positioning relative to labels
  • Table structure and column alignment
  • Presence of signatures, stamps, or annotations
  • Document formatting and visual hierarchy

Future Evolution

Current multimodal capabilities are maturing rapidly. othman-moumni-abdou predicts that eventually models will be accurate enough with just image input, potentially making the OCR transcription step unnecessary as pure visual processing capabilities improve.

Production Implementation

Requirements

  • Multimodal LLM capability (text + image input)
  • OCR pipeline for text extraction
  • Image preprocessing and formatting
  • Validation framework for both modalities

Performance Monitoring

Production systems need evaluation frameworks that can assess:

  • Text extraction accuracy vs image interpretation accuracy
  • Multimodal combination effectiveness
  • Regression detection when updating either OCR or vision components

See also

Production RAG Systems

page dédiée →

Operational considerations and quality standards required for deploying Retrieval-Augmented Generation systems in production environments, including security, reliability, and maintainability requirements.

Security Requirements

Input Validation

All user inputs must be validated and sanitized:

  • Metadata filters must use parameterized queries to prevent sql-injection-in-llm-systems
  • Document uploads require content type validation and virus scanning
  • Query parameters need bounds checking and encoding validation

Access Controls

  • Authentication for system access
  • Authorization for document collections
  • Audit logging for compliance tracking
  • Rate limiting to prevent abuse

Reliability Standards

Error Handling

Production systems require comprehensive error handling:

# Graceful degradation when retrieval fails
try:
    results = retriever.search(query)
except DatabaseError:
    logger.error("Retrieval failed, using fallback")
    results = fallback_search(query)

Dependencies Management

  • Import validation at startup to catch missing dependencies
  • Dependency pinning for reproducible deployments
  • Health checks for external services
  • Circuit breakers for service failures

Performance Monitoring

  • Response time tracking for user experience
  • Throughput monitoring for capacity planning
  • Resource utilization monitoring
  • Cache hit rates for optimization

Code Quality Standards

Module Structure

  • Clean imports with proper dependency management
  • Constructor consistency across components
  • Parameter validation in public interfaces
  • Type annotations for maintainability

Configuration Management

  • Environment-based configuration for different deployments
  • Configuration validation at startup
  • Secret management for API keys and credentials
  • Feature flags for safe rollouts

Common Pitfalls

Based on analysis of assistant-rh and other systems:

Security Issues

  • Direct string interpolation in database queries
  • Unvalidated file uploads and processing
  • Missing authentication on admin endpoints
  • Insufficient logging of security events

Implementation Problems

  • Missing or incorrect imports preventing module loading
  • Constructor argument mismatches causing runtime failures
  • Unhandled exceptions causing service crashes
  • Resource leaks in file processing

Architectural Issues

  • Tight coupling between components making testing difficult
  • Lack of graceful degradation when services fail
  • Missing health check endpoints for monitoring
  • Poor error propagation and logging

Deployment Checklist

Before production deployment:

  • Security audit including penetration testing
  • Load testing to validate performance under load
  • Disaster recovery procedures documented and tested
  • Monitoring and alerting configured for critical paths
  • Documentation updated for operational procedures
  • Rollback procedures tested and validated

See also

Production Systems Evolution

page dédiée →

The maturation of LLM-based production systems from initial automation achievements to sophisticated evaluation and quality control frameworks. Demonstrated through alan-health's document processing evolution, showing how real-world deployment drives system architecture improvements and operational practices.

Evolution Stages

Initial Deployment (2024-2025)

  • Focus on basic automation and accuracy metrics
  • Text-only processing approaches
  • Simple success/failure evaluation
  • Manual quality control processes

Mature Production (2025-2026)

Advanced Operations

  • Field-level regression tracking with criticality weights
  • approximate-nearest-neighbor-search for scalable example selection
  • Separation of pure parsing from enrichment processes
  • Cross-industry pattern validation

Key Learning Areas

System Architecture Maturity

Production deployment reveals architectural bottlenecks not apparent in development:

  • classification-bottleneck as single point of failure
  • Need for robust validation frameworks
  • Importance of modular, testable components

Operational Practices Evolution

Real-world usage drives operational sophistication:

  • "Measure before you ship" principle
  • Structured error handling and human review routing
  • Continuous improvement through validated document accumulation

Quality Control Development

Production quality requirements drive evaluation framework sophistication:

  • Reference dataset curation and management
  • Multi-dimensional accuracy tracking
  • Regression prevention mechanisms

Production Insights Pattern

Challenge Identification

Running systems at scale reveals problems invisible in lab settings:

  • document-quality-challenges in real-world data
  • Few-shot example contamination effects
  • Classification accuracy impact on downstream processing

Solution Development

Production constraints drive practical solution development:

  • Hybrid approaches balancing accuracy and performance
  • Scalable evaluation methods for continuous deployment
  • Error recovery and human-in-the-loop integration

Knowledge Generalization

Production learnings establish industry patterns:

  • Cross-validation between healthcare and financial sectors
  • Transferable architectural principles
  • Reusable operational frameworks

Industry Impact

Knowledge Transfer

Production insights from pioneers like othman-moumni-abdou accelerate industry maturation by documenting real-world challenges and solutions.

Standard Practices Emergence

Repeated patterns across organizations establish industry best practices:

  • Multimodal input superiority
  • Layout-based similarity matching
  • Comprehensive evaluation frameworks

Technology Evolution Driver

Production requirements drive technology advancement:

  • Model capability improvements
  • Tool and framework development
  • Infrastructure optimization

See also

Pure Parsing

page dédiée →

Document extraction approach that limits LLM output to information directly visible in the source document, excluding external knowledge, cross-document context, or inferred information. Critical architectural pattern developed at alan-health to prevent few-shot-contamination and maintain extraction reliability in production systems.

Core Principle

Document-Grounded Extraction

Constraint: Extract only information explicitly present in the source document

  • Visible text: Information that can be read directly from document content
  • Visual elements: Data represented in images, tables, stamps, signatures
  • No inference: Avoid filling missing information from domain knowledge
  • No context: Exclude information from other documents or external sources

Separation from Enrichment

Two-phase architecture: Distinct separation between extraction and business logic

  1. Pure parsing phase: Document → structured data (document-grounded only)
  2. Post-processing phase: Structured data + business rules → enriched output

Problem: External Knowledge Contamination

Human Operator Behavior

Manual document processors naturally apply external knowledge:

  • Cross-document lookups: Checking related documents in same claim
  • Domain expertise: Applying healthcare industry knowledge
  • External validation: Confirming procedure codes against standard references
  • Business rules: Filling missing fields based on organizational policies

Training Data Corruption

When human-processed examples become few-shot training data:

  • Contaminated examples: Training data includes non-document information
  • Hallucination learning: LLM learns to invent plausible missing data
  • Extraction drift: Model output gradually includes more external knowledge
  • Audit trail loss: Cannot verify extracted data against source documents

Solution Architecture

Pure Parsing Implementation

Extraction constraints enforced through:

  • Prompt engineering: Explicit instructions to extract only visible information
  • Example curation: Few-shot datasets verified to contain only document-grounded data
  • Validation rules: Schema checks ensuring extracted fields map to document content
  • Human training: Manual processors educated on pure parsing principles

Post-Processing Enrichment

Business logic applied separately:

  • Cross-document context: Integration with related documents after pure extraction
  • External lookups: API calls to validation services and reference databases
  • Domain knowledge: Application of business rules and healthcare expertise
  • Inference logic: Filling missing fields based on organizational requirements

Production Benefits

Extraction Reliability

  • Hallucination prevention: Eliminates fabricated data not grounded in documents
  • Audit transparency: Clear mapping between extracted data and source content
  • Quality consistency: Predictable extraction behavior independent of operator knowledge
  • Training stability: Few-shot examples remain document-grounded over time

System Modularity

  • Independent evolution: Extraction logic can improve without affecting business rules
  • Business rule flexibility: Post-processing can change without retraining extraction
  • Clear responsibility: Distinct ownership of document processing vs business enrichment
  • Testing isolation: Extraction accuracy can be measured independently

Implementation Challenges

Prompt Engineering

Clarity requirements: Instructions must clearly distinguish visible vs inferred information

  • Negative examples: Show what NOT to extract from external knowledge
  • Boundary cases: Handle situations where document content is ambiguous
  • Validation rules: Define exactly what constitutes "visible" information

Example Quality Control

Reference dataset curation: Systematic removal of contaminated training examples

  • Human review: Verify examples contain only document-visible information
  • Audit trails: Maintain source tracking for all training examples
  • Continuous cleaning: Regular review and updating of reference datasets

Performance Impact

Information loss: Some useful enrichment must wait for post-processing phase

  • Field completeness: Pure extraction may leave more fields empty
  • Processing complexity: Two-phase architecture adds system complexity
  • Performance tradeoffs: Multiple processing steps vs single enriched extraction

Validation Techniques

Document Grounding Verification

  • Source highlighting: Verify extracted fields can be highlighted in original document
  • Human validation: Manual spot-checks of extraction vs document content
  • Automated checks: Schema validation ensuring extracted data types match document structure
  • Comparative testing: Pure parsing accuracy vs contaminated extraction

See also

Pure Parsing Separation

page dédiée →

Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.

Problem Addressed

When human operators create training examples, they often include information not visible in the source document:

  • Cross-document context from related claims
  • External knowledge (procedure codes, standard prices)
  • Domain expertise and business logic application

Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.

Architectural Solution

Two-Stage Processing

  1. Pure Parsing Stage: LLM extracts only what's directly visible in the document
  2. Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context

Benefits

  • Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
  • Reduces Hallucinations: LLM cannot invent values based on external context
  • Improves Accuracy: Document extraction becomes more reliable and predictable
  • Enables Validation: Pure parsing results can be verified against source documents

Implementation Requirements

Clean Training Data

  • Review existing few-shot examples for external contamination
  • Separate document-visible information from inferred information
  • Create pure parsing examples that contain only document content

Architectural Separation

  • Independent pure parsing component focused solely on document content
  • Separate post-processing pipeline for enrichment and business logic
  • Clear interface between parsing and enrichment stages

Validation Framework

  • Verify pure parsing outputs against source documents
  • Monitor for hallucination patterns indicating contamination
  • Test enrichment logic independently of parsing accuracy

Production Benefits

alan-health identified this separation as crucial for production reliability:

  • More predictable LLM behavior in document extraction
  • Easier debugging when issues arise (parsing vs enrichment problems)
  • Better evaluation capability for each pipeline stage
  • Reduced maintenance overhead from hallucination issues

See also

Pydantic Validation

page dédiée →

Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.

Core Concept

Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.

Basic Implementation Pattern

from pydantic import BaseModel
from typing import Literal

class LLMResponse(BaseModel):
    content: str
    confidence: float
    action_type: Literal["question", "answer", "guess"]
    
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)

Production Applications

Document Processing Pipeline (alan-health)

  • Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
  • Human Review Integration: Structured rejection flows when validation fails
  • Audit Trail: Maintains validation history for compliance requirements
  • Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review

Game Logic (lesphinx)

  • Deterministic Behavior: SphinxAction model ensures AI responses contain required fields
  • Action Classification: action_type field enables branching logic (question vs. guess)
  • Confidence Tracking: Validated confidence scores for game flow control
  • Multilingual Support: Schema validation across language boundaries

Key Validation Patterns

Field-Level Constraints

class DocumentExtraction(BaseModel):
    medical_code: str = Field(min_length=3, max_length=10)
    confidence: float = Field(ge=0.0, le=1.0)  # Between 0 and 1
    review_required: bool

Enum Validation for Controlled Vocabularies

class GameAction(BaseModel):
    action_type: Literal["question", "guess", "end"]
    language: Literal["fr", "en"]

Custom Validators

from pydantic import validator

class MedicalDocument(BaseModel):
    @validator('medical_code')
    def validate_code_format(cls, v):
        if not v.startswith('ICD'):
            raise ValueError('Medical code must start with ICD')
        return v

Error Handling Strategies

Validation Failure Recovery

  1. Immediate Retry: Re-prompt LLM with validation error context
  2. Fallback Schemas: Simpler models when complex validation fails
  3. Human Handoff: Queue for manual review when automated validation consistently fails
  4. Default Values: Safe defaults for non-critical fields

JSON Security

Critical for preventing injection attacks in LLM-generated content:

# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'

# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})

Production Benefits

Reliability

  • Type Safety: Compile-time checking prevents runtime type errors
  • Data Integrity: Ensures downstream systems receive expected data formats
  • Graceful Degradation: Structured error handling when LLM outputs are malformed

Maintainability

  • Schema Evolution: Versioned models enable backward-compatible API changes
  • Documentation: Pydantic models serve as living documentation of data contracts
  • Testing: Easy mock generation for unit tests using model factories

Observability

  • Validation Metrics: Track validation success rates across different LLM providers
  • Error Classification: Categorize validation failures for model improvement
  • Performance Monitoring: Measure validation overhead in processing pipelines

Integration with LLM APIs

Structured Generation

Many LLM providers support JSON mode or schema-guided generation:

# OpenAI JSON mode
response = openai.chat.completions.create(
    model="gpt-4",
    response_format={"type": "json_object"},
    messages=[...]
)

# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)

Response Parsing Pipeline

def process_llm_response(raw_response: str, schema: Type[BaseModel]):
    try:
        # Primary validation attempt
        return schema.model_validate_json(raw_response)
    except ValidationError as e:
        # Log validation error details
        logger.error(f"Validation failed: {e}")
        # Attempt retry or fallback
        return handle_validation_failure(raw_response, schema, e)

Anti-Patterns

Over-Validation

  • Excessive Constraints: Too strict validation can cause high failure rates
  • Rigid Schemas: Inflexible models that break with reasonable LLM variations

Under-Validation

  • Missing Required Fields: Optional fields that should be required for business logic
  • Weak Type Constraints: Using Any or str when specific types are needed

Error Handling Gaps

  • Silent Failures: Catching validation errors without proper logging or recovery
  • Blocking Operations: Synchronous validation that blocks processing pipelines

See also