Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Classification Bottleneck
page dédiée →Critical single point of failure in document processing pipelines where incorrect document classification leads to wrong extraction schema application, making results unusable. Major production challenge identified at alan-health processing French healthcare documents, where misclassification renders subsequent extraction steps ineffective regardless of extraction quality.
Problem Definition
Single Point of Failure: If classifier predicts wrong document category, extraction step applies incorrect schema to document, producing unusable structured output.
Chain Reaction: Even perfect extraction logic fails when operating on wrong document type assumptions - attempting to extract invoice fields from prescription document structure.
Quality Amplification: Classification errors amplify downstream, turning high-quality extraction capabilities into garbage output due to schema mismatch.
Specific Failure Modes
Vocabulary Overlap: French hospital attestations frequently misclassified as emergency invoices due to shared medical terminology and similar document structure patterns.
Concatenated Documents: Users upload multiple document types in single PDF (prescription + invoice + payment receipt), confusing classifiers designed to expect one document type per upload.
Subtle Distinctions: Similar document layouts with different business purposes require nuanced classification that current systems struggle with at production scale.
Production Impact at Scale
alan-health identifies classification accuracy as limiting factor for overall pipeline performance:
- Perfect extraction becomes worthless with wrong classification
- Human review required for misclassified documents regardless of extraction confidence
- Single bottleneck preventing automation rate improvements
Mitigation Strategies
Multi-Modal Classification: Leverage both text content and visual layout for improved category distinction.
Confidence Thresholding: Route low-confidence classifications to human review before extraction attempt.
Document Splitting: Detect and handle concatenated documents through layout analysis or content segmentation.
Active Learning: Continuously improve classifier on production misclassification examples.
See also
Concatenated Document Challenges
page dédiée →Classification problem in document processing where users upload multiple document types combined into a single PDF file, confusing classifiers designed to expect one document type per upload. Major challenge identified at alan-health causing classification failures and downstream extraction errors.
Problem Description
User Behavior: Members upload concatenated documents containing prescription, invoice, and payment receipt all in single PDF file for convenience.
Classifier Confusion: Document classification systems designed to identify single document type per file fail when presented with multiple document types in sequence.
Schema Mismatch: Single-type extraction schemas cannot handle multi-document content, leading to extraction failure or incorrect field mapping.
Classification Impact
Single Point of Failure: Concatenated documents trigger classification-bottleneck where wrong or unclear classification makes entire document processing pipeline unusable.
Schema Selection Failure: No appropriate extraction schema exists for multi-document content, forcing fallback to human review.
Automation Rate Reduction: High-frequency concatenated uploads reduce overall automation rates despite individual document type processing working correctly.
Technical Challenges
Layout Analysis: Detecting document boundaries within concatenated PDF requires sophisticated layout analysis beyond simple page breaks.
Context Switching: Extraction logic must recognize when document type changes mid-stream and apply appropriate schema transitions.
Validation Complexity: Combined documents require validation logic that handles multiple schema types within single processing pipeline.
Mitigation Strategies
Document Splitting Detection: Implement layout analysis to identify distinct document boundaries within concatenated files.
Multi-Stage Processing: Process detected document segments separately with appropriate classification and extraction for each segment.
User Education: Guide users toward single-document uploads while maintaining fallback handling for concatenated cases.
Preprocessing Pipeline: Automatic document separation preprocessing step before classification and extraction.
Production Context
alan-health actively working to improve classification accuracy for concatenated document edge cases as part of broader classification-bottleneck resolution efforts.
See also
Concatenated Document Classification
page dédiée →Challenge in document processing where users upload multiple document types in a single PDF, confusing classifiers designed to expect one document type per upload. Major production issue identified at alan-health processing French healthcare documents.
Problem Definition
Typical Concatenated Patterns
- Prescription + invoice + payment receipt in single PDF
- Hospital attestation + multiple invoices
- Mixed healthcare document types spanning multiple pages
Classification System Confusion
- Traditional classifiers trained on single-document-type assumptions
- Vocabulary overlap between document types creates ambiguity
- First document in sequence may drive classification decision
- Extraction schema mismatch leads to unusable results
Impact on Processing Pipeline
Single Point of Failure
- Incorrect classification renders entire extraction unusable
- Wrong schema applied to mixed content produces garbled results
- Human review required even for otherwise processable individual documents
Example: French Healthcare Documents
- Hospital attestations misclassified as emergency invoices due to shared vocabulary
- Prescription documents combined with pharmacy invoices confuse categorical boundaries
- Payment receipts concatenated with medical invoices create classification uncertainty
Potential Solutions
Document Segmentation
- Pre-processing to identify document boundaries within PDF
- Individual classification per document segment
- Parallel extraction pipelines for each identified document type
Multi-Label Classification
- Update classification system to handle multiple document types
- Confidence scoring per document type detected
- Sequential processing based on detected types
Layout-Based Detection
- Visual layout analysis to identify document boundaries
- Page-level classification before content-level analysis
- Structural cues for document type identification
Current Status
alan-health identifies this as active area for improvement in their production system. Classification accuracy improvements needed for these edge cases to reduce human review requirements and improve automation rates.
See also
Configurable Pipeline Design
page dédiée →Architectural approach to document processing pipelines that enables deployment across multiple countries and use cases without fundamental system changes. Upcoming focus area for alan-health's document processing system, building on their successful French healthcare document automation.
Core Concept
Rather than building separate pipelines for each market or document type, create a configurable system that can adapt to:
- Different country-specific document formats
- Varying regulatory requirements
- Multiple languages and scripts
- Distinct healthcare system structures
- Alternative business logic requirements
Strategic Importance
As alan-health expands beyond French markets, the ability to quickly configure their 70% automation pipeline for new countries becomes critical for scaling operations efficiently.
Implementation Considerations
- Schema flexibility for country-specific fields
- Language-agnostic OCR integration
- Configurable validation rules
- Market-specific few-shot example datasets
- Adaptable classification categories
Future Development
othman-moumni-abdou has indicated this will be covered in an upcoming article in Alan's document processing series, detailing specific implementation strategies and lessons learned.
See also
- document-processing-pipeline
- alan-health
- othman-moumni-abdou
Cross-Industry Document Processing
page dédiée →Validation that production LLM document processing techniques and challenges generalize across industries, with patterns identified in healthcare applying equally to financial, legal, and other structured document domains. Key evidence comes from holofin's confirmation of alan-health's insights.
Universal Production Patterns
Common Challenges Across Industries:
- Few-Shot Contamination: External knowledge bleeding into examples affects healthcare, financial, and legal document processing equally
- Classification Bottlenecks: Single points of failure in document categorization impact all structured document workflows
- Document Quality Issues: Handwriting, poor scans, and image degradation affect all industries processing physical documents
- Layout-Based Similarity: Visual structure matching outperforms semantic content matching across domains
Cross-Industry Validation Examples
Healthcare → Financial: edouard-foussier at holofin confirmed that layout-based similarity matching, originally developed for French healthcare documents at alan-health, applies equally to financial document processing.
Technical Principles: Multimodal approaches, pure parsing separation, and evaluation frameworks show consistent value across healthcare and financial document automation.
Why Techniques Generalize
Structural Similarity: Business documents across industries follow similar patterns:
- Form-based layouts with predictable field arrangements
- Table structures requiring spatial understanding
- Template-based designs with visual hierarchy
- Mixed content types (text, numbers, signatures, stamps)
LLM Behavior Consistency: Core challenges with LLM document processing stem from model characteristics rather than domain specifics:
- Hallucination tendencies when examples contain external knowledge
- Superior performance with combined text+image inputs
- Need for systematic evaluation and quality control
Implementation Benefits
Shared Best Practices: Organizations can adopt proven techniques from other industries rather than reinventing solutions.
Accelerated Development: Cross-industry validation reduces experimentation time and risk when building new document processing systems.
Universal Architecture: Core pipeline components (classification, extraction, validation, evaluation) apply broadly with domain-specific customization only in business logic layers.
Industry-Specific Adaptations
While core techniques generalize, each industry requires customization in:
- Schema Design: Field types and validation rules specific to document types
- Regulatory Compliance: GDPR for healthcare, SOX for financial, industry-specific requirements
- Domain Knowledge: Post-processing enrichment using industry-specific databases and rules
- Quality Standards: Criticality weights and accuracy thresholds based on business impact
See also
- layout-based-similarity-matching
- few-shot-contamination
- alan-health
- holofin
- edouard-foussier
- Production LLM Systems
- document-quality-challenges
Cross-Industry Validation
page dédiée →The process of confirming that production AI techniques and insights developed in one industry apply effectively to other sectors, establishing general principles rather than domain-specific solutions. Critical for determining the broader applicability and transferability of production AI practices.
Key Validation Example
Healthcare to Financial Document Processing
othman-moumni-abdou's production insights from alan-health processing French healthcare documents received validation from edouard-foussier at holofin working on financial documents. This cross-industry confirmation established several techniques as broadly applicable:
- Layout-Based Similarity: Visual structure matching for few-shot examples works across both healthcare and financial documents
- Multimodal Processing: OCR + image input advantages apply beyond healthcare
- Few-Shot Contamination Challenges: Human-enriched examples create similar problems across industries
- Production Evaluation Needs: Systematic backtesting requirements generalize
Significance for Production AI
Technique Generalizability
Cross-industry validation increases confidence that specific production techniques represent fundamental principles rather than domain-specific optimizations. This allows:
- Faster adoption across industries
- Higher confidence in implementation decisions
- Reduced need for domain-specific research
- Better resource allocation for technique development
Pattern Recognition
Validation across industries helps identify which challenges are universal vs domain-specific:
- Universal: Document quality issues, classification bottlenecks, evaluation framework needs
- Domain-Specific: Vocabulary overlap patterns, regulatory requirements, specific document types
Implementation Value
Risk Reduction
Cross-industry validation provides evidence that production techniques will work in new contexts, reducing implementation risk and development time.
Knowledge Transfer
Successful patterns from one industry can be rapidly adapted to others, accelerating production AI development across sectors.
Future Applications
This validation pattern suggests other production AI insights from healthcare, finance, legal, and other document-intensive industries may have broader applicability than initially recognized.
See also
Curated Reference Datasets
page dédiée →Small, hand-picked collections of high-quality document examples with verified extractions used to bootstrap new document processing categories. Key insight from alan-health: a small, carefully curated reference dataset enables rapid deployment of new document types without requiring large validated pools.
Core Principle
Quality over Quantity: A small set of hand-picked reference documents with verified extractions gets you "surprisingly far" when bootstrapping new categories. Focus on representative, clean examples rather than large volumes of potentially inconsistent data.
Implementation Approach
Manual Curation: Human experts select representative examples covering typical document variations within a category.
Verified Ground Truth: Each reference document has manually validated, high-quality extraction serving as ground truth for evaluation and few-shot selection.
Category Bootstrapping: New document types can be launched with minimal initial data, then improved iteratively as more validated examples accumulate.
Production Benefits
- Enables rapid deployment of new document processing categories
- Provides reliable foundation for few-shot-learning without large data requirements
- Supports evaluation-framework-design with trusted ground truth references
- Scales naturally as system processes more documents over time
See also
Document Processing Pipeline
page dédiée →Production-grade system for extracting structured data from documents using LLMs, particularly effective for complex healthcare and insurance documents. Combines OCR transcription with document images for optimal accuracy, as demonstrated by alan-health's 70% automation rate on French healthcare documents.
Architecture Evolution
Text-Only Processing (Initial Approach)
- Input: OCR Markdown transcription only
- Output: Structured data extraction
- Performance: Good baseline accuracy for most documents
Image-Only Processing (Experimental)
- Input: Document images only, no OCR
- Findings: Lower accuracy than text-based extraction
- Problems: Hallucinations where LLM "read" information not actually on documents
- Conclusion: OCR text provides more reliable parsing foundation than raw pixel interpretation
Multimodal Processing (Current Best Practice)
- Input: OCR Markdown transcription + document images combined
- Performance: Outperforms either input alone
- Benefits:
- OCR text provides reliable, precise content parsing
- Images provide visual layout context, field positioning, table structure
- Combined approach leverages strengths of both modalities
Production Components
Validation Layer
- Pydantic schema validation for structured output
- Failed validation triggers human review with structured error messages
- Best-effort extraction preserved as starting point for reviewers
Human-in-the-Loop Integration
- Configurable human review for high-stakes document categories
- Systematic review for new document categories during rollout
- Online evaluation comparison with manual parsing results
Reference Dataset Strategy
- Curated examples bootstrap new document categories
- Small, hand-picked datasets achieve surprising effectiveness
- Validated documents become potential few-shot examples over time
Future Evolution
Current multimodal capabilities are maturing rapidly. As models improve image-only processing accuracy, the OCR transcription step may become unnecessary, simplifying the pipeline architecture.
See also
Few-Shot Contamination
page dédiée →Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.
Problem Description
Root Cause: Human operators processing documents don't limit themselves to visible content. They:
- Check other documents in same claim for context
- Look up healthcare procedure codes and standard prices online
- Apply domain knowledge not present in document text
- Make inferences from external systems or databases
Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.
Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."
Production Impact
Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.
Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.
Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.
Solution: Architectural Separation
Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.
Post-Processing Phase: Separate step for enrichment using:
- Business logic and rules
- Cross-document context lookup
- External API calls and database queries
- Domain knowledge application
Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.
Implementation at Scale
alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.
See also
- document-processing-pipeline
- Reference Dataset Design
- LLM Hallucination
- Production AI Systems
- alan-health
HNSW Indexing
page dédiée →Hierarchical Navigable Small World (HNSW) indexing technique used for approximate nearest neighbor search in document processing pipelines, trading small precision loss for significant speed improvements when searching large document example pools.
Problem Context
For high-volume document categories with millions of reference examples, exact L2 distance search becomes computationally expensive. alan-health implemented HNSW indexing to maintain performance as their reference document pools grew to massive scale.
Technical Approach
Exact vs Approximate Search
- Exact L2 Distance: Accurate but slower as example pools grow
- HNSW Approximate: Small precision trade-off for major speed improvements
- Use Case: High-volume document categories requiring fast similarity matching
Implementation Benefits
- Maintains acceptable accuracy for few-shot example selection
- Scales to millions of reference documents
- Enables real-time similarity matching in production
- Supports efficient nearest neighbor queries for document matching
Production Context
Part of alan-health's evolved document processing pipeline, used alongside:
- curated-reference-datasets for smaller document categories
- layout-based-similarity matching approaches
- multimodal-llm-processing for combined text + image input
Performance Characteristics
HNSW provides a configurable trade-off between:
- Search speed (higher with HNSW)
- Search precision (slightly lower with HNSW)
- Memory efficiency
- Scalability to large document pools
See also
- approximate-nearest-neighbor-search
- Document Processing Performance
- Similarity Search Optimization
Layout-Based Similarity
page dédiée →Document processing technique that matches few-shot examples based on visual layout structure rather than semantic text content, improving extraction accuracy by providing structurally similar reference documents. Key insight from edouard-foussier at holofin, validated across both financial and healthcare document processing.
Core Principle
Instead of selecting few-shot examples based on textual similarity or semantic content overlap, this approach prioritizes documents with similar visual structures:
- Table layouts and positioning
- Field arrangement and spacing
- Header/footer structures
- Visual organization patterns
Cross-Industry Validation
Originally identified as effective for financial documents at holofin, this technique validates insights from alan-health's healthcare document processing pipeline. The consistency across industries suggests this is a fundamental principle for multimodal document processing rather than domain-specific optimization.
Implementation Benefits
Improved Extraction Accuracy
Visual layout similarity provides better context for LLM extraction because:
- Document structure strongly correlates with field positioning
- Similar layouts indicate similar extraction patterns
- Visual cues guide field identification more effectively than text similarity
Reduced Few-Shot Contamination
By focusing on layout rather than content, this approach helps avoid semantic bias that can lead to hallucinated values, addressing the few-shot contamination problem identified in production systems.
Technical Implementation
Requires capability to:
- Analyze document visual structure
- Create layout-based similarity metrics
- Select examples based on structural matching rather than content matching
- Potentially use computer vision techniques for layout analysis
Production Impact
edouard-foussier's confirmation at holofin validates this as a production-ready technique that improves document processing accuracy across different document types and industries, establishing it as a best practice for multimodal document processing systems.
See also
Multimodal Document Processing
page dédiée →Production approach combining both OCR text transcription and document images as input to LLM systems, outperforming either single-modal approach. Represents evolved best practice as of 2025-2026, demonstrated at scale by alan-health processing French healthcare documents.
Evolution Path
Text-Only → Image-Only → Multimodal
-
Text-Only Processing: Initial approach using only OCR Markdown transcription as LLM input. Worked well for most documents but lacked visual context.
-
Image-Only Experiment: Attempted processing with document images alone, skipping transcription. Result: Lower accuracy than text-based extraction, with increased hallucinations where LLMs "read" information not actually present in ambiguous or hard-to-parse images.
-
Multimodal Combination: Current best practice combining both OCR transcription and document images. Outperforms either input alone.
Why Multimodal Works
Complementary Information Sources:
- OCR Transcription Provides: Reliable text content that LLMs can parse precisely, reducing ambiguity in character recognition
- Document Images Provide: Visual layout context, spatial relationships, table structures, presence of stamps/signatures, form organization
Reduced Hallucination Risk: OCR text anchors the extraction to actual document content, while images provide visual verification and context that prevents misinterpretation.
Technical Implementation
Input Format: Send both OCR Markdown transcription AND document image to multimodal LLM models simultaneously.
Processing Strategy: LLM uses text for precise content extraction while referencing image for layout understanding and visual verification.
Production Benefits:
- Higher extraction accuracy than single-modal approaches
- Better handling of complex table structures
- Improved recognition of visual elements (stamps, signatures)
- Reduced hallucination on ambiguous content
Future Evolution
As of early 2025-2026, multimodal LLMs still require OCR transcription support for optimal accuracy. However, visual capabilities are maturing rapidly - at some point, models may achieve sufficient accuracy with image-only input, potentially eliminating the OCR transcription step.
Production Validation
alan-health's production system processes millions of French healthcare documents using this multimodal approach, achieving 70% automation rates with higher accuracy than previous single-modal systems.
See also
- document-quality-challenges
- OCR Limitations
- Visual Layout Context
- othman-moumni-abdou
- alan-health
- Production LLM Systems
Multimodal LLM Processing
page dédiée →Advanced document processing approach combining both OCR-extracted text and document images as input to LLMs, achieving superior extraction accuracy compared to either modality alone. Key insight from alan-health: text provides reliable content parsing while images provide crucial visual layout context.
Processing Evolution
Single Modality Limitations
Text-Only Processing: Initial approach using only OCR Markdown transcription worked well for most documents but missed visual context cues.
Image-Only Processing: Experimental approach achieved lower accuracy than text-based extraction, with significant hallucination problems where LLMs would "read" information not actually present in ambiguous or hard-to-parse images.
Optimal Multimodal Approach
Combined OCR + Image: Current best practice providing superior results through:
- OCR Text: Reliable content that LLMs can parse precisely
- Document Image: Visual layout context showing field positioning, table structures, stamps, signatures
Technical Advantages
Complementary Information
- Text provides precise character-level content extraction
- Images provide spatial relationships and visual formatting context
- Combined input reduces ambiguity in document structure interpretation
Hallucination Reduction
Image-only processing produced hallucinations where LLMs invented values from ambiguous visual content. OCR text anchors the LLM to actual document content while images provide confirmatory visual context.
Layout Context
Visual elements crucial for accurate extraction:
- Field positioning relative to labels
- Table structure and column alignment
- Presence of signatures, stamps, or annotations
- Document formatting and visual hierarchy
Future Evolution
Current multimodal capabilities are maturing rapidly. othman-moumni-abdou predicts that eventually models will be accurate enough with just image input, potentially making the OCR transcription step unnecessary as pure visual processing capabilities improve.
Production Implementation
Requirements
- Multimodal LLM capability (text + image input)
- OCR pipeline for text extraction
- Image preprocessing and formatting
- Validation framework for both modalities
Performance Monitoring
Production systems need evaluation frameworks that can assess:
- Text extraction accuracy vs image interpretation accuracy
- Multimodal combination effectiveness
- Regression detection when updating either OCR or vision components
See also
Production AI Hiring
page dédiée →Focus area for companies building production LLM systems, requiring engineers who can handle the unique challenges of making AI work reliably on messy real-world data at scale. alan-health actively recruiting for engineers interested in document processing, LLM reliability, and production AI systems.
Required Skills
- Document processing pipeline development
- LLM reliability engineering
- Real-world data handling at scale
- Production system architecture
- Evaluation framework design
Challenge Areas
- Making AI work on messy, real-world data
- Scaling LLM systems to high-volume production
- Building reliable evaluation and monitoring systems
- Handling edge cases and data quality issues
- Designing human-in-the-loop workflows
Market Demand
The production AI engineering field requires specialized expertise that combines traditional software engineering with deep understanding of LLM behavior, evaluation methodologies, and the unique challenges of deploying AI systems in enterprise environments.
See also
Pure Parsing Separation
page dédiée →Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.
Problem Addressed
When human operators create training examples, they often include information not visible in the source document:
- Cross-document context from related claims
- External knowledge (procedure codes, standard prices)
- Domain expertise and business logic application
Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.
Architectural Solution
Two-Stage Processing
- Pure Parsing Stage: LLM extracts only what's directly visible in the document
- Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context
Benefits
- Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
- Reduces Hallucinations: LLM cannot invent values based on external context
- Improves Accuracy: Document extraction becomes more reliable and predictable
- Enables Validation: Pure parsing results can be verified against source documents
Implementation Requirements
Clean Training Data
- Review existing few-shot examples for external contamination
- Separate document-visible information from inferred information
- Create pure parsing examples that contain only document content
Architectural Separation
- Independent pure parsing component focused solely on document content
- Separate post-processing pipeline for enrichment and business logic
- Clear interface between parsing and enrichment stages
Validation Framework
- Verify pure parsing outputs against source documents
- Monitor for hallucination patterns indicating contamination
- Test enrichment logic independently of parsing accuracy
Production Benefits
alan-health identified this separation as crucial for production reliability:
- More predictable LLM behavior in document extraction
- Easier debugging when issues arise (parsing vs enrichment problems)
- Better evaluation capability for each pipeline stage
- Reduced maintenance overhead from hallucination issues
See also
- few-shot-contamination
- Document Processing Architecture
- LLM Hallucination Prevention
Pydantic Validation
page dédiée →Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.
Core Concept
Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.
Basic Implementation Pattern
from pydantic import BaseModel
from typing import Literal
class LLMResponse(BaseModel):
content: str
confidence: float
action_type: Literal["question", "answer", "guess"]
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)
Production Applications
Document Processing Pipeline (alan-health)
- Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
- Human Review Integration: Structured rejection flows when validation fails
- Audit Trail: Maintains validation history for compliance requirements
- Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review
Game Logic (lesphinx)
- Deterministic Behavior:
SphinxActionmodel ensures AI responses contain required fields - Action Classification:
action_typefield enables branching logic (question vs. guess) - Confidence Tracking: Validated confidence scores for game flow control
- Multilingual Support: Schema validation across language boundaries
Key Validation Patterns
Field-Level Constraints
class DocumentExtraction(BaseModel):
medical_code: str = Field(min_length=3, max_length=10)
confidence: float = Field(ge=0.0, le=1.0) # Between 0 and 1
review_required: bool
Enum Validation for Controlled Vocabularies
class GameAction(BaseModel):
action_type: Literal["question", "guess", "end"]
language: Literal["fr", "en"]
Custom Validators
from pydantic import validator
class MedicalDocument(BaseModel):
@validator('medical_code')
def validate_code_format(cls, v):
if not v.startswith('ICD'):
raise ValueError('Medical code must start with ICD')
return v
Error Handling Strategies
Validation Failure Recovery
- Immediate Retry: Re-prompt LLM with validation error context
- Fallback Schemas: Simpler models when complex validation fails
- Human Handoff: Queue for manual review when automated validation consistently fails
- Default Values: Safe defaults for non-critical fields
JSON Security
Critical for preventing injection attacks in LLM-generated content:
# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'
# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})
Production Benefits
Reliability
- Type Safety: Compile-time checking prevents runtime type errors
- Data Integrity: Ensures downstream systems receive expected data formats
- Graceful Degradation: Structured error handling when LLM outputs are malformed
Maintainability
- Schema Evolution: Versioned models enable backward-compatible API changes
- Documentation: Pydantic models serve as living documentation of data contracts
- Testing: Easy mock generation for unit tests using model factories
Observability
- Validation Metrics: Track validation success rates across different LLM providers
- Error Classification: Categorize validation failures for model improvement
- Performance Monitoring: Measure validation overhead in processing pipelines
Integration with LLM APIs
Structured Generation
Many LLM providers support JSON mode or schema-guided generation:
# OpenAI JSON mode
response = openai.chat.completions.create(
model="gpt-4",
response_format={"type": "json_object"},
messages=[...]
)
# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)
Response Parsing Pipeline
def process_llm_response(raw_response: str, schema: Type[BaseModel]):
try:
# Primary validation attempt
return schema.model_validate_json(raw_response)
except ValidationError as e:
# Log validation error details
logger.error(f"Validation failed: {e}")
# Attempt retry or fallback
return handle_validation_failure(raw_response, schema, e)
Anti-Patterns
Over-Validation
- Excessive Constraints: Too strict validation can cause high failure rates
- Rigid Schemas: Inflexible models that break with reasonable LLM variations
Under-Validation
- Missing Required Fields: Optional fields that should be required for business logic
- Weak Type Constraints: Using
Anyorstrwhen specific types are needed
Error Handling Gaps
- Silent Failures: Catching validation errors without proper logging or recovery
- Blocking Operations: Synchronous validation that blocks processing pipelines
See also
- llm-integration-patterns
- structured-output
- json-security
- error-handling
- alan-health
- lesphinx
RAG vs Wiki Pattern
page dédiée →Fundamental comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern, highlighting different approaches to knowledge management and query answering.
Traditional RAG Approach
Process Flow
- Upload collection of files
- Chunk and embed documents
- At query time: retrieve relevant chunks
- Generate answer from retrieved fragments
- Repeat process for each new query
Characteristics
- Stateless: Each query starts fresh
- Rediscovery: Knowledge reconstructed every time
- Fragment-based: Works with document chunks, not integrated knowledge
- No accumulation: Understanding doesn't build over time
- Limited synthesis: Difficult to connect insights across multiple documents
Examples
- NotebookLM
- ChatGPT file uploads
- Most commercial RAG systems
- Document Q&A tools
LLM Wiki Pattern Approach
Process Flow
- Ingest sources into persistent wiki structure
- LLM builds and maintains cross-referenced knowledge base
- At query time: synthesize from existing wiki pages
- File valuable answers back as new wiki content
- Knowledge compounds with each interaction
Characteristics
- Stateful: Persistent knowledge accumulation
- Pre-synthesized: Understanding built incrementally over time
- Integration-based: Works with structured, cross-referenced content
- Compounding: Each addition makes the whole more valuable
- Rich synthesis: Connections already established across sources
Key Differences
| Aspect | Traditional RAG | Wiki Pattern |
|---|---|---|
| Knowledge State | Stateless retrieval | Persistent accumulation |
| Processing Time | Query-time discovery | Ingest-time integration |
| Cross-reference | Ad-hoc during query | Pre-established and maintained |
| Contradictions | Discovered per query | Flagged and tracked systematically |
| Synthesis Quality | Limited by retrieval | Rich, pre-built connections |
| Maintenance | None (static chunks) | Automated wiki maintenance |
When to Use Each
RAG Works Well For
- Simple document Q&A
- One-off queries against large document sets
- When you don't need knowledge to compound
- Rapid deployment without setup overhead
- Documents that rarely need cross-referencing
Wiki Pattern Works Well For
- Long-term knowledge building
- Complex synthesis across multiple sources
- Research that builds over time
- When contradictions and evolution matter
- Domains where connections between concepts are valuable
Hybrid Approaches
Some systems might combine both patterns:
- Wiki pattern for core, frequently-accessed knowledge
- RAG for supplementary document collections
- Different layers for different types of content
Performance Implications
RAG
- Pros: Simple setup, works immediately, scales to large document sets
- Cons: Repeated processing overhead, limited synthesis depth, no knowledge accumulation
Wiki Pattern
- Pros: Rich synthesis, compounding value, deep cross-references, maintained consistency
- Cons: Higher setup cost, requires ongoing LLM maintenance, more complex architecture