Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Backtest Monitoring
page dédiée →Production evaluation system that re-runs document processing pipelines on reference datasets with verified ground truth to measure the impact of changes before deployment. Essential safety mechanism implemented at alan-health to prevent silent regressions across document categories.
Core Functionality
Pipeline Re-execution: Complete processing workflow from document input through transcription, classification, and extraction steps on curated reference datasets.
Ground Truth Comparison: Field-by-field comparison between system output and verified expected results across entire extraction schema.
Change Impact Measurement: Quantifies improvements and regressions before changes reach production environment.
Monitoring Components
Classification Diff Analysis:
- Document category prediction accuracy
- Sub-class classification performance
- Category-specific error patterns
Extraction Diff Analysis:
- Field-level accuracy comparison
- Value-by-value extraction validation
- Schema compliance measurement
Criticality Weighting System:
- High priority: Financial amounts, critical dates, regulatory fields
- Medium priority: Names, addresses, secondary identifiers
- Low priority: Formatting variations, optional fields
Dashboard Integration
Visual Monitoring: Real-time tracking of accuracy metrics across document categories with regression alerts and improvement validation.
Aggregate Analysis: Results analyzed per category, per field, or in aggregate to identify patterns and guide development decisions.
Historical Tracking: Longitudinal performance monitoring enabling trend analysis and regression root cause identification.
Production Implementation
Pre-Deployment Validation: Every system change validated against reference datasets before production release.
Automated Safety Net: Prevents deployment of changes that degrade performance on any document category.
Development Confidence: Teams iterate with quantified impact visibility rather than subjective assessment.
Reference Dataset Requirements
Production Authenticity: Real documents from production pipeline with verified manual extractions as ground truth.
Representative Coverage: Spans all document types, quality levels, and edge cases encountered in production environment.
Immutable Standards: Stable reference datasets ensure consistent measurement across evaluation runs.
See also
- evaluation-framework-design
- Reference Dataset Design
- Production AI Systems
- alan-health
- document-processing-pipeline
Clean Architecture
page dédiée →Software design philosophy emphasizing separation of concerns, dependency inversion, and autonomous components to create maintainable, testable, and evolution-friendly systems. Particularly valuable for production systems requiring long-term maintainability and team handovers.
Core Principles
Separation of Concerns
Each module should have a single, well-defined responsibility:
# Instead of monolithic pipeline.py (1561 lines)
pipeline.py # Orchestration only (300 lines)
query_processor.py # Query analysis and reformulation
retriever.py # Multi-source data retrieval
section_aggregator.py # Chunk to section transformation
context_builder.py # Document aggregation logic
generator.py # Response generation
Dependency Inversion
High-level modules should not depend on low-level modules. Both should depend on abstractions:
# Bad: Direct dependency on specific implementation
from src.rag_v2.embedder import ScalewayEmbedder
# Good: Internal abstraction
from .embedder import Embedder # Internalized interface
Autonomous Components
Each major component should be self-contained with minimal external dependencies. This enables independent testing, deployment, and evolution.
Implementation Strategies
Layered Architecture
Presentation Layer: UI components, API endpoints, user interfaces
Application Layer: Business logic, orchestration, workflow management
Domain Layer: Core business entities, rules, and domain-specific logic
Infrastructure Layer: Database access, external APIs, system interfaces
Component Internalization
Rather than maintaining complex shared libraries, internalize and simplify components for specific use cases:
# Before: General-purpose shared component
class FallbackEmbedder:
"""Supports 5 embedding providers with complex fallback logic"""
# 425 lines handling multiple providers, retry logic, caching, etc.
# After: Clean, specific implementation
class Embedder:
"""Albert + Scaleway embeddings for HR document processing"""
# 200 lines focused on actual requirements
Interface Consistency
Define clear contracts between components that remain stable even as implementations change:
@dataclass
class RetrievalResult:
chunks: List[DocumentChunk]
sections: List[DocumentSection]
confidence: float
source_tables: List[str]
class Retriever:
def retrieve(self, query: str, top_k: int = 20) -> RetrievalResult:
"""Consistent interface regardless of implementation details"""
Real-World Application: RAG System Migration
The assistant-rh clean architecture migration demonstrates these principles in practice:
Before (Unclear Separation):
pipeline.py: 1561 lines mixing orchestration, business logic, and data access- Complex imports across
Concatenated Document Classification
page dédiée →Challenge in document processing where users upload multiple document types in a single PDF, confusing classifiers designed to expect one document type per upload. Major production issue identified at alan-health processing French healthcare documents.
Problem Definition
Typical Concatenated Patterns
- Prescription + invoice + payment receipt in single PDF
- Hospital attestation + multiple invoices
- Mixed healthcare document types spanning multiple pages
Classification System Confusion
- Traditional classifiers trained on single-document-type assumptions
- Vocabulary overlap between document types creates ambiguity
- First document in sequence may drive classification decision
- Extraction schema mismatch leads to unusable results
Impact on Processing Pipeline
Single Point of Failure
- Incorrect classification renders entire extraction unusable
- Wrong schema applied to mixed content produces garbled results
- Human review required even for otherwise processable individual documents
Example: French Healthcare Documents
- Hospital attestations misclassified as emergency invoices due to shared vocabulary
- Prescription documents combined with pharmacy invoices confuse categorical boundaries
- Payment receipts concatenated with medical invoices create classification uncertainty
Potential Solutions
Document Segmentation
- Pre-processing to identify document boundaries within PDF
- Individual classification per document segment
- Parallel extraction pipelines for each identified document type
Multi-Label Classification
- Update classification system to handle multiple document types
- Confidence scoring per document type detected
- Sequential processing based on detected types
Layout-Based Detection
- Visual layout analysis to identify document boundaries
- Page-level classification before content-level analysis
- Structural cues for document type identification
Current Status
alan-health identifies this as active area for improvement in their production system. Classification accuracy improvements needed for these edge cases to reduce human review requirements and improve automation rates.
See also
Document Processing Pipeline
page dédiée →Production-grade system for extracting structured data from documents using LLMs, particularly effective for complex healthcare and insurance documents. Combines OCR transcription with document images for optimal accuracy, as demonstrated by alan-health's 70% automation rate on French healthcare documents.
Architecture Evolution
Text-Only Processing (Initial Approach)
- Input: OCR Markdown transcription only
- Output: Structured data extraction
- Performance: Good baseline accuracy for most documents
Image-Only Processing (Experimental)
- Input: Document images only, no OCR
- Findings: Lower accuracy than text-based extraction
- Problems: Hallucinations where LLM "read" information not actually on documents
- Conclusion: OCR text provides more reliable parsing foundation than raw pixel interpretation
Multimodal Processing (Current Best Practice)
- Input: OCR Markdown transcription + document images combined
- Performance: Outperforms either input alone
- Benefits:
- OCR text provides reliable, precise content parsing
- Images provide visual layout context, field positioning, table structure
- Combined approach leverages strengths of both modalities
Production Components
Validation Layer
- Pydantic schema validation for structured output
- Failed validation triggers human review with structured error messages
- Best-effort extraction preserved as starting point for reviewers
Human-in-the-Loop Integration
- Configurable human review for high-stakes document categories
- Systematic review for new document categories during rollout
- Online evaluation comparison with manual parsing results
Reference Dataset Strategy
- Curated examples bootstrap new document categories
- Small, hand-picked datasets achieve surprising effectiveness
- Validated documents become potential few-shot examples over time
Future Evolution
Current multimodal capabilities are maturing rapidly. As models improve image-only processing accuracy, the OCR transcription step may become unnecessary, simplifying the pipeline architecture.
See also
Evaluation Framework Design
page dédiée →Systematic approach for measuring LLM system changes before production deployment, preventing silent regressions and ensuring quality improvements. Essential for production systems where changes must be validated against real-world performance, as implemented at alan-health with comprehensive backtest monitoring and field-level analysis.
Core Principle: Measure Before You Ship
Pre-Deployment Validation: Every change (new LLM model, updated instructions, different few-shot strategy) measured against reference datasets before reaching production.
Regression Prevention: Avoid scenarios where pharmacy invoice extraction improvements silently degrade hospital invoice accuracy.
Confidence Building: Teams can iterate on extraction instructions, switch models, or adjust strategies with quantified impact visibility.
Reference Dataset Requirements
Ground Truth Quality: Real production documents with verified extractions serving as gold standard for comparison.
Representative Coverage: Documents spanning all categories, edge cases, and quality levels encountered in production.
Versioned and Immutable: Reference datasets remain stable for consistent measurement across evaluation runs.
Evaluation Components
Classification Diff: Measure whether system predicts correct document category and sub-classes compared to expected results.
Extraction Diff: Field-by-field comparison between actual extracted values and expected results across entire schema.
Criticality Weighting: Different importance levels for extraction errors:
- High criticality: Amount paid, dates, critical financial values
- Medium criticality: Healthcare professional names, secondary fields
- Low criticality: Minor formatting or optional field variations
Backtest Execution
Full Pipeline Re-run: Complete processing from document input through transcription, classification, and extraction steps.
Automated Analysis: Systematic comparison storing results for analysis per category, per field, or in aggregate.
Dashboard Monitoring: Visual tracking of accuracy metrics, regression alerts, and improvement validation across document types.
Production Implementation
alan-health uses this framework for:
- LLM model upgrades and downgrades
- Prompt engineering iterations
- Few-shot strategy modifications
- Pipeline architecture changes
- Performance optimization validation
Results stored and analyzed to guide development decisions with quantified confidence rather than subjective assessment.
See also
- backtest-monitoring
- Reference Dataset Design
- Production AI Systems
- alan-health
- document-processing-pipeline
Few-Shot Contamination
page dédiée →Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.
Problem Description
Root Cause: Human operators processing documents don't limit themselves to visible content. They:
- Check other documents in same claim for context
- Look up healthcare procedure codes and standard prices online
- Apply domain knowledge not present in document text
- Make inferences from external systems or databases
Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.
Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."
Production Impact
Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.
Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.
Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.
Solution: Architectural Separation
Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.
Post-Processing Phase: Separate step for enrichment using:
- Business logic and rules
- Cross-document context lookup
- External API calls and database queries
- Domain knowledge application
Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.
Implementation at Scale
alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.
See also
- document-processing-pipeline
- Reference Dataset Design
- LLM Hallucination
- Production AI Systems
- alan-health
Interface Consistency
page dédiée →Fundamental software architecture principle requiring that components implementing the same interface return compatible data types and follow identical behavioral contracts. Critical for system reliability and maintainability, particularly in complex systems like RAG pipelines where multiple implementations must be interchangeable.
Core Principle
Interface consistency ensures that:
- All implementations of an interface return the same data types
- Method signatures are identical across implementations
- Behavioral contracts are maintained regardless of underlying implementation
- Downstream consumers can rely on predictable data formats
Critical Failure Mode: Mixed Return Types
A common violation occurs when different implementations of the same interface return incompatible data types:
# Inconsistent interface implementation
class TFIDFRetriever:
def search(self, query):
return [(index, score), (index, score)] # Returns tuples
class PostgresRetriever:
def search(self, query):
return [Chunk(...), Chunk(...)] # Returns objects
This creates downstream failures when consumers expect consistent types:
# Downstream consumer crashes with tuples
def build_prompt(chunks):
for chunk in chunks:
content = chunk.content # AttributeError if chunk is tuple
Production Impact
Interface inconsistency can cause:
- Runtime Crashes: Type errors when unexpected data types are processed
- Silent Data Corruption: Partial processing with wrong assumptions
- Debugging Complexity: Failures occur far from the source of inconsistency
- System Brittleness: Adding new implementations breaks existing code
RAG System Implications
In rag-systems, interface consistency is particularly critical because:
Multiple Retriever Implementations
Systems often support multiple retrieval strategies:
- Vector-based retrieval returning ranked documents
- Keyword-based retrieval returning scored matches
- Hybrid systems combining multiple approaches
- multi-source-rag federating across different backends
Downstream Processing Chains
Retrieved results flow through multiple processing stages:
- Reranking systems expecting specific document formats
- Context builders assembling prompts from document content
- Logging systems recording search results
- UI components displaying search results
Pluggable Architecture Requirements
Production RAG systems require:
- Hot-swappable retriever implementations
- A/B testing with different retrieval strategies
- Fallback mechanisms when primary retrievers fail
- Configuration-driven retriever selection
Design Patterns for Consistency
Common Interface Definition
from abc import ABC, abstractmethod
from typing import List
class Retriever(ABC):
@abstractmethod
def search(self, query: str, limit: int = 10) -> List[Chunk]:
"""Return list of Chunk objects ranked by relevance"""
pass
Adapter Pattern for Legacy Systems
class LegacyRetrieverAdapter(Retriever):
def __init__(self, legacy_retriever):
self.legacy = legacy_retriever
def search(self, query: str, limit: int = 10) -> List[Chunk]:
# Convert legacy tuple format to Chunk objects
raw_results = self.legacy.search(query, limit)
return [self._tuple_to_chunk(r) for r in raw_results]
Runtime Type Validation
def validate_search_results(results: List[Chunk]) -> List[Chunk]:
"""Validate that all results are proper Chunk objects"""
for result in results:
if not isinstance(result, Chunk):
raise TypeError(f"Expected Chunk, got {type(result)}")
return results
Testing Strategies
Interface Compliance Tests
def test_retriever_interface_compliance():
retrievers = [TFIDFRetriever(), PostgresRetriever(), HybridRetriever()]
for retriever in retrievers:
results = retriever.search("test query")
# Verify return type consistency
assert isinstance(results, list)
for result in results:
assert isinstance(result, Chunk)
assert hasattr(result, 'content')
assert hasattr(result, 'score')
Downstream Integration Tests
def test_build_prompt_with_all_retrievers():
"""Test that build_prompt works with results from any retriever"""
test_query = "sample query"
for retriever in all_retriever_implementations():
results = retriever.search(test_query)
prompt = build_prompt(results) # Should not crash
assert isinstance(prompt, str)
assert len(prompt) > 0
See also
- rag-systems
- software-architecture
- multi-source-rag
- contract-programming
- system-reliability
Multimodal Document Processing
page dédiée →Production approach combining both OCR text transcription and document images as input to LLM systems, outperforming either single-modal approach. Represents evolved best practice as of 2025-2026, demonstrated at scale by alan-health processing French healthcare documents.
Evolution Path
Text-Only → Image-Only → Multimodal
-
Text-Only Processing: Initial approach using only OCR Markdown transcription as LLM input. Worked well for most documents but lacked visual context.
-
Image-Only Experiment: Attempted processing with document images alone, skipping transcription. Result: Lower accuracy than text-based extraction, with increased hallucinations where LLMs "read" information not actually present in ambiguous or hard-to-parse images.
-
Multimodal Combination: Current best practice combining both OCR transcription and document images. Outperforms either input alone.
Why Multimodal Works
Complementary Information Sources:
- OCR Transcription Provides: Reliable text content that LLMs can parse precisely, reducing ambiguity in character recognition
- Document Images Provide: Visual layout context, spatial relationships, table structures, presence of stamps/signatures, form organization
Reduced Hallucination Risk: OCR text anchors the extraction to actual document content, while images provide visual verification and context that prevents misinterpretation.
Technical Implementation
Input Format: Send both OCR Markdown transcription AND document image to multimodal LLM models simultaneously.
Processing Strategy: LLM uses text for precise content extraction while referencing image for layout understanding and visual verification.
Production Benefits:
- Higher extraction accuracy than single-modal approaches
- Better handling of complex table structures
- Improved recognition of visual elements (stamps, signatures)
- Reduced hallucination on ambiguous content
Future Evolution
As of early 2025-2026, multimodal LLMs still require OCR transcription support for optimal accuracy. However, visual capabilities are maturing rapidly - at some point, models may achieve sufficient accuracy with image-only input, potentially eliminating the OCR transcription step.
Production Validation
alan-health's production system processes millions of French healthcare documents using this multimodal approach, achieving 70% automation rates with higher accuracy than previous single-modal systems.
See also
- document-quality-challenges
- OCR Limitations
- Visual Layout Context
- othman-moumni-abdou
- alan-health
- Production LLM Systems
Multimodal LLM Processing
page dédiée →Advanced document processing approach combining both OCR-extracted text and document images as input to LLMs, achieving superior extraction accuracy compared to either modality alone. Key insight from alan-health: text provides reliable content parsing while images provide crucial visual layout context.
Processing Evolution
Single Modality Limitations
Text-Only Processing: Initial approach using only OCR Markdown transcription worked well for most documents but missed visual context cues.
Image-Only Processing: Experimental approach achieved lower accuracy than text-based extraction, with significant hallucination problems where LLMs would "read" information not actually present in ambiguous or hard-to-parse images.
Optimal Multimodal Approach
Combined OCR + Image: Current best practice providing superior results through:
- OCR Text: Reliable content that LLMs can parse precisely
- Document Image: Visual layout context showing field positioning, table structures, stamps, signatures
Technical Advantages
Complementary Information
- Text provides precise character-level content extraction
- Images provide spatial relationships and visual formatting context
- Combined input reduces ambiguity in document structure interpretation
Hallucination Reduction
Image-only processing produced hallucinations where LLMs invented values from ambiguous visual content. OCR text anchors the LLM to actual document content while images provide confirmatory visual context.
Layout Context
Visual elements crucial for accurate extraction:
- Field positioning relative to labels
- Table structure and column alignment
- Presence of signatures, stamps, or annotations
- Document formatting and visual hierarchy
Future Evolution
Current multimodal capabilities are maturing rapidly. othman-moumni-abdou predicts that eventually models will be accurate enough with just image input, potentially making the OCR transcription step unnecessary as pure visual processing capabilities improve.
Production Implementation
Requirements
- Multimodal LLM capability (text + image input)
- OCR pipeline for text extraction
- Image preprocessing and formatting
- Validation framework for both modalities
Performance Monitoring
Production systems need evaluation frameworks that can assess:
- Text extraction accuracy vs image interpretation accuracy
- Multimodal combination effectiveness
- Regression detection when updating either OCR or vision components
See also
Production RAG Systems
page dédiée →Operational considerations and quality standards required for deploying Retrieval-Augmented Generation systems in production environments, including security, reliability, and maintainability requirements.
Security Requirements
Input Validation
All user inputs must be validated and sanitized:
- Metadata filters must use parameterized queries to prevent sql-injection-in-llm-systems
- Document uploads require content type validation and virus scanning
- Query parameters need bounds checking and encoding validation
Access Controls
- Authentication for system access
- Authorization for document collections
- Audit logging for compliance tracking
- Rate limiting to prevent abuse
Reliability Standards
Error Handling
Production systems require comprehensive error handling:
# Graceful degradation when retrieval fails
try:
results = retriever.search(query)
except DatabaseError:
logger.error("Retrieval failed, using fallback")
results = fallback_search(query)
Dependencies Management
- Import validation at startup to catch missing dependencies
- Dependency pinning for reproducible deployments
- Health checks for external services
- Circuit breakers for service failures
Performance Monitoring
- Response time tracking for user experience
- Throughput monitoring for capacity planning
- Resource utilization monitoring
- Cache hit rates for optimization
Code Quality Standards
Module Structure
- Clean imports with proper dependency management
- Constructor consistency across components
- Parameter validation in public interfaces
- Type annotations for maintainability
Configuration Management
- Environment-based configuration for different deployments
- Configuration validation at startup
- Secret management for API keys and credentials
- Feature flags for safe rollouts
Common Pitfalls
Based on analysis of assistant-rh and other systems:
Security Issues
- Direct string interpolation in database queries
- Unvalidated file uploads and processing
- Missing authentication on admin endpoints
- Insufficient logging of security events
Implementation Problems
- Missing or incorrect imports preventing module loading
- Constructor argument mismatches causing runtime failures
- Unhandled exceptions causing service crashes
- Resource leaks in file processing
Architectural Issues
- Tight coupling between components making testing difficult
- Lack of graceful degradation when services fail
- Missing health check endpoints for monitoring
- Poor error propagation and logging
Deployment Checklist
Before production deployment:
- Security audit including penetration testing
- Load testing to validate performance under load
- Disaster recovery procedures documented and tested
- Monitoring and alerting configured for critical paths
- Documentation updated for operational procedures
- Rollback procedures tested and validated
See also
- sql-injection-in-llm-systems
- rag-pipeline-architecture
- automated-evaluation
- assistant-rh - Case study in production readiness issues
Production Systems Evolution
page dédiée →The maturation of LLM-based production systems from initial automation achievements to sophisticated evaluation and quality control frameworks. Demonstrated through alan-health's document processing evolution, showing how real-world deployment drives system architecture improvements and operational practices.
Evolution Stages
Initial Deployment (2024-2025)
- Focus on basic automation and accuracy metrics
- Text-only processing approaches
- Simple success/failure evaluation
- Manual quality control processes
Mature Production (2025-2026)
- multimodal-llm-processing with combined text+image input
- Comprehensive evaluation-framework-design with backtest monitoring
- curated-reference-datasets for rapid category bootstrapping
- Systematic few-shot-contamination prevention
Advanced Operations
- Field-level regression tracking with criticality weights
- approximate-nearest-neighbor-search for scalable example selection
- Separation of pure parsing from enrichment processes
- Cross-industry pattern validation
Key Learning Areas
System Architecture Maturity
Production deployment reveals architectural bottlenecks not apparent in development:
- classification-bottleneck as single point of failure
- Need for robust validation frameworks
- Importance of modular, testable components
Operational Practices Evolution
Real-world usage drives operational sophistication:
- "Measure before you ship" principle
- Structured error handling and human review routing
- Continuous improvement through validated document accumulation
Quality Control Development
Production quality requirements drive evaluation framework sophistication:
- Reference dataset curation and management
- Multi-dimensional accuracy tracking
- Regression prevention mechanisms
Production Insights Pattern
Challenge Identification
Running systems at scale reveals problems invisible in lab settings:
- document-quality-challenges in real-world data
- Few-shot example contamination effects
- Classification accuracy impact on downstream processing
Solution Development
Production constraints drive practical solution development:
- Hybrid approaches balancing accuracy and performance
- Scalable evaluation methods for continuous deployment
- Error recovery and human-in-the-loop integration
Knowledge Generalization
Production learnings establish industry patterns:
- Cross-validation between healthcare and financial sectors
- Transferable architectural principles
- Reusable operational frameworks
Industry Impact
Knowledge Transfer
Production insights from pioneers like othman-moumni-abdou accelerate industry maturation by documenting real-world challenges and solutions.
Standard Practices Emergence
Repeated patterns across organizations establish industry best practices:
- Multimodal input superiority
- Layout-based similarity matching
- Comprehensive evaluation frameworks
Technology Evolution Driver
Production requirements drive technology advancement:
- Model capability improvements
- Tool and framework development
- Infrastructure optimization
See also
- alan-health - Organization demonstrating this evolution
- othman-moumni-abdou - Researcher documenting evolution patterns
- evaluation-framework-design - Key component of system maturity
- multimodal-llm-processing - Technical evolution example
- cross-industry-validation - Pattern generalization across sectors
Pure Parsing
page dédiée →Document extraction approach that limits LLM output to information directly visible in the source document, excluding external knowledge, cross-document context, or inferred information. Critical architectural pattern developed at alan-health to prevent few-shot-contamination and maintain extraction reliability in production systems.
Core Principle
Document-Grounded Extraction
Constraint: Extract only information explicitly present in the source document
- Visible text: Information that can be read directly from document content
- Visual elements: Data represented in images, tables, stamps, signatures
- No inference: Avoid filling missing information from domain knowledge
- No context: Exclude information from other documents or external sources
Separation from Enrichment
Two-phase architecture: Distinct separation between extraction and business logic
- Pure parsing phase: Document → structured data (document-grounded only)
- Post-processing phase: Structured data + business rules → enriched output
Problem: External Knowledge Contamination
Human Operator Behavior
Manual document processors naturally apply external knowledge:
- Cross-document lookups: Checking related documents in same claim
- Domain expertise: Applying healthcare industry knowledge
- External validation: Confirming procedure codes against standard references
- Business rules: Filling missing fields based on organizational policies
Training Data Corruption
When human-processed examples become few-shot training data:
- Contaminated examples: Training data includes non-document information
- Hallucination learning: LLM learns to invent plausible missing data
- Extraction drift: Model output gradually includes more external knowledge
- Audit trail loss: Cannot verify extracted data against source documents
Solution Architecture
Pure Parsing Implementation
Extraction constraints enforced through:
- Prompt engineering: Explicit instructions to extract only visible information
- Example curation: Few-shot datasets verified to contain only document-grounded data
- Validation rules: Schema checks ensuring extracted fields map to document content
- Human training: Manual processors educated on pure parsing principles
Post-Processing Enrichment
Business logic applied separately:
- Cross-document context: Integration with related documents after pure extraction
- External lookups: API calls to validation services and reference databases
- Domain knowledge: Application of business rules and healthcare expertise
- Inference logic: Filling missing fields based on organizational requirements
Production Benefits
Extraction Reliability
- Hallucination prevention: Eliminates fabricated data not grounded in documents
- Audit transparency: Clear mapping between extracted data and source content
- Quality consistency: Predictable extraction behavior independent of operator knowledge
- Training stability: Few-shot examples remain document-grounded over time
System Modularity
- Independent evolution: Extraction logic can improve without affecting business rules
- Business rule flexibility: Post-processing can change without retraining extraction
- Clear responsibility: Distinct ownership of document processing vs business enrichment
- Testing isolation: Extraction accuracy can be measured independently
Implementation Challenges
Prompt Engineering
Clarity requirements: Instructions must clearly distinguish visible vs inferred information
- Negative examples: Show what NOT to extract from external knowledge
- Boundary cases: Handle situations where document content is ambiguous
- Validation rules: Define exactly what constitutes "visible" information
Example Quality Control
Reference dataset curation: Systematic removal of contaminated training examples
- Human review: Verify examples contain only document-visible information
- Audit trails: Maintain source tracking for all training examples
- Continuous cleaning: Regular review and updating of reference datasets
Performance Impact
Information loss: Some useful enrichment must wait for post-processing phase
- Field completeness: Pure extraction may leave more fields empty
- Processing complexity: Two-phase architecture adds system complexity
- Performance tradeoffs: Multiple processing steps vs single enriched extraction
Validation Techniques
Document Grounding Verification
- Source highlighting: Verify extracted fields can be highlighted in original document
- Human validation: Manual spot-checks of extraction vs document content
- Automated checks: Schema validation ensuring extracted data types match document structure
- Comparative testing: Pure parsing accuracy vs contaminated extraction
See also
Pure Parsing Separation
page dédiée →Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.
Problem Addressed
When human operators create training examples, they often include information not visible in the source document:
- Cross-document context from related claims
- External knowledge (procedure codes, standard prices)
- Domain expertise and business logic application
Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.
Architectural Solution
Two-Stage Processing
- Pure Parsing Stage: LLM extracts only what's directly visible in the document
- Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context
Benefits
- Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
- Reduces Hallucinations: LLM cannot invent values based on external context
- Improves Accuracy: Document extraction becomes more reliable and predictable
- Enables Validation: Pure parsing results can be verified against source documents
Implementation Requirements
Clean Training Data
- Review existing few-shot examples for external contamination
- Separate document-visible information from inferred information
- Create pure parsing examples that contain only document content
Architectural Separation
- Independent pure parsing component focused solely on document content
- Separate post-processing pipeline for enrichment and business logic
- Clear interface between parsing and enrichment stages
Validation Framework
- Verify pure parsing outputs against source documents
- Monitor for hallucination patterns indicating contamination
- Test enrichment logic independently of parsing accuracy
Production Benefits
alan-health identified this separation as crucial for production reliability:
- More predictable LLM behavior in document extraction
- Easier debugging when issues arise (parsing vs enrichment problems)
- Better evaluation capability for each pipeline stage
- Reduced maintenance overhead from hallucination issues
See also
- few-shot-contamination
- Document Processing Architecture
- LLM Hallucination Prevention
Pydantic Validation
page dédiée →Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.
Core Concept
Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.
Basic Implementation Pattern
from pydantic import BaseModel
from typing import Literal
class LLMResponse(BaseModel):
content: str
confidence: float
action_type: Literal["question", "answer", "guess"]
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)
Production Applications
Document Processing Pipeline (alan-health)
- Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
- Human Review Integration: Structured rejection flows when validation fails
- Audit Trail: Maintains validation history for compliance requirements
- Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review
Game Logic (lesphinx)
- Deterministic Behavior:
SphinxActionmodel ensures AI responses contain required fields - Action Classification:
action_typefield enables branching logic (question vs. guess) - Confidence Tracking: Validated confidence scores for game flow control
- Multilingual Support: Schema validation across language boundaries
Key Validation Patterns
Field-Level Constraints
class DocumentExtraction(BaseModel):
medical_code: str = Field(min_length=3, max_length=10)
confidence: float = Field(ge=0.0, le=1.0) # Between 0 and 1
review_required: bool
Enum Validation for Controlled Vocabularies
class GameAction(BaseModel):
action_type: Literal["question", "guess", "end"]
language: Literal["fr", "en"]
Custom Validators
from pydantic import validator
class MedicalDocument(BaseModel):
@validator('medical_code')
def validate_code_format(cls, v):
if not v.startswith('ICD'):
raise ValueError('Medical code must start with ICD')
return v
Error Handling Strategies
Validation Failure Recovery
- Immediate Retry: Re-prompt LLM with validation error context
- Fallback Schemas: Simpler models when complex validation fails
- Human Handoff: Queue for manual review when automated validation consistently fails
- Default Values: Safe defaults for non-critical fields
JSON Security
Critical for preventing injection attacks in LLM-generated content:
# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'
# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})
Production Benefits
Reliability
- Type Safety: Compile-time checking prevents runtime type errors
- Data Integrity: Ensures downstream systems receive expected data formats
- Graceful Degradation: Structured error handling when LLM outputs are malformed
Maintainability
- Schema Evolution: Versioned models enable backward-compatible API changes
- Documentation: Pydantic models serve as living documentation of data contracts
- Testing: Easy mock generation for unit tests using model factories
Observability
- Validation Metrics: Track validation success rates across different LLM providers
- Error Classification: Categorize validation failures for model improvement
- Performance Monitoring: Measure validation overhead in processing pipelines
Integration with LLM APIs
Structured Generation
Many LLM providers support JSON mode or schema-guided generation:
# OpenAI JSON mode
response = openai.chat.completions.create(
model="gpt-4",
response_format={"type": "json_object"},
messages=[...]
)
# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)
Response Parsing Pipeline
def process_llm_response(raw_response: str, schema: Type[BaseModel]):
try:
# Primary validation attempt
return schema.model_validate_json(raw_response)
except ValidationError as e:
# Log validation error details
logger.error(f"Validation failed: {e}")
# Attempt retry or fallback
return handle_validation_failure(raw_response, schema, e)
Anti-Patterns
Over-Validation
- Excessive Constraints: Too strict validation can cause high failure rates
- Rigid Schemas: Inflexible models that break with reasonable LLM variations
Under-Validation
- Missing Required Fields: Optional fields that should be required for business logic
- Weak Type Constraints: Using
Anyorstrwhen specific types are needed
Error Handling Gaps
- Silent Failures: Catching validation errors without proper logging or recovery
- Blocking Operations: Synchronous validation that blocks processing pipelines
See also
- llm-integration-patterns
- structured-output
- json-security
- error-handling
- alan-health
- lesphinx