~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Classification Bottleneck

page dédiée →

Critical single point of failure in document processing pipelines where incorrect document classification leads to wrong extraction schema application, making results unusable. Major production challenge identified at alan-health processing French healthcare documents, where misclassification renders subsequent extraction steps ineffective regardless of extraction quality.

Problem Definition

Single Point of Failure: If classifier predicts wrong document category, extraction step applies incorrect schema to document, producing unusable structured output.

Chain Reaction: Even perfect extraction logic fails when operating on wrong document type assumptions - attempting to extract invoice fields from prescription document structure.

Quality Amplification: Classification errors amplify downstream, turning high-quality extraction capabilities into garbage output due to schema mismatch.

Specific Failure Modes

Vocabulary Overlap: French hospital attestations frequently misclassified as emergency invoices due to shared medical terminology and similar document structure patterns.

Concatenated Documents: Users upload multiple document types in single PDF (prescription + invoice + payment receipt), confusing classifiers designed to expect one document type per upload.

Subtle Distinctions: Similar document layouts with different business purposes require nuanced classification that current systems struggle with at production scale.

Production Impact at Scale

alan-health identifies classification accuracy as limiting factor for overall pipeline performance:

  • Perfect extraction becomes worthless with wrong classification
  • Human review required for misclassified documents regardless of extraction confidence
  • Single bottleneck preventing automation rate improvements

Mitigation Strategies

Multi-Modal Classification: Leverage both text content and visual layout for improved category distinction.

Confidence Thresholding: Route low-confidence classifications to human review before extraction attempt.

Document Splitting: Detect and handle concatenated documents through layout analysis or content segmentation.

Active Learning: Continuously improve classifier on production misclassification examples.

See also

Concatenated Document Challenges

page dédiée →

Classification problem in document processing where users upload multiple document types combined into a single PDF file, confusing classifiers designed to expect one document type per upload. Major challenge identified at alan-health causing classification failures and downstream extraction errors.

Problem Description

User Behavior: Members upload concatenated documents containing prescription, invoice, and payment receipt all in single PDF file for convenience.

Classifier Confusion: Document classification systems designed to identify single document type per file fail when presented with multiple document types in sequence.

Schema Mismatch: Single-type extraction schemas cannot handle multi-document content, leading to extraction failure or incorrect field mapping.

Classification Impact

Single Point of Failure: Concatenated documents trigger classification-bottleneck where wrong or unclear classification makes entire document processing pipeline unusable.

Schema Selection Failure: No appropriate extraction schema exists for multi-document content, forcing fallback to human review.

Automation Rate Reduction: High-frequency concatenated uploads reduce overall automation rates despite individual document type processing working correctly.

Technical Challenges

Layout Analysis: Detecting document boundaries within concatenated PDF requires sophisticated layout analysis beyond simple page breaks.

Context Switching: Extraction logic must recognize when document type changes mid-stream and apply appropriate schema transitions.

Validation Complexity: Combined documents require validation logic that handles multiple schema types within single processing pipeline.

Mitigation Strategies

Document Splitting Detection: Implement layout analysis to identify distinct document boundaries within concatenated files.

Multi-Stage Processing: Process detected document segments separately with appropriate classification and extraction for each segment.

User Education: Guide users toward single-document uploads while maintaining fallback handling for concatenated cases.

Preprocessing Pipeline: Automatic document separation preprocessing step before classification and extraction.

Production Context

alan-health actively working to improve classification accuracy for concatenated document edge cases as part of broader classification-bottleneck resolution efforts.

See also

Concatenated Document Classification

page dédiée →

Challenge in document processing where users upload multiple document types in a single PDF, confusing classifiers designed to expect one document type per upload. Major production issue identified at alan-health processing French healthcare documents.

Problem Definition

Typical Concatenated Patterns

  • Prescription + invoice + payment receipt in single PDF
  • Hospital attestation + multiple invoices
  • Mixed healthcare document types spanning multiple pages

Classification System Confusion

  • Traditional classifiers trained on single-document-type assumptions
  • Vocabulary overlap between document types creates ambiguity
  • First document in sequence may drive classification decision
  • Extraction schema mismatch leads to unusable results

Impact on Processing Pipeline

Single Point of Failure

  • Incorrect classification renders entire extraction unusable
  • Wrong schema applied to mixed content produces garbled results
  • Human review required even for otherwise processable individual documents

Example: French Healthcare Documents

  • Hospital attestations misclassified as emergency invoices due to shared vocabulary
  • Prescription documents combined with pharmacy invoices confuse categorical boundaries
  • Payment receipts concatenated with medical invoices create classification uncertainty

Potential Solutions

Document Segmentation

  • Pre-processing to identify document boundaries within PDF
  • Individual classification per document segment
  • Parallel extraction pipelines for each identified document type

Multi-Label Classification

  • Update classification system to handle multiple document types
  • Confidence scoring per document type detected
  • Sequential processing based on detected types

Layout-Based Detection

  • Visual layout analysis to identify document boundaries
  • Page-level classification before content-level analysis
  • Structural cues for document type identification

Current Status

alan-health identifies this as active area for improvement in their production system. Classification accuracy improvements needed for these edge cases to reduce human review requirements and improve automation rates.

See also

Configurable Pipeline Design

page dédiée →

Architectural approach to document processing pipelines that enables deployment across multiple countries and use cases without fundamental system changes. Upcoming focus area for alan-health's document processing system, building on their successful French healthcare document automation.

Core Concept

Rather than building separate pipelines for each market or document type, create a configurable system that can adapt to:

  • Different country-specific document formats
  • Varying regulatory requirements
  • Multiple languages and scripts
  • Distinct healthcare system structures
  • Alternative business logic requirements

Strategic Importance

As alan-health expands beyond French markets, the ability to quickly configure their 70% automation pipeline for new countries becomes critical for scaling operations efficiently.

Implementation Considerations

  • Schema flexibility for country-specific fields
  • Language-agnostic OCR integration
  • Configurable validation rules
  • Market-specific few-shot example datasets
  • Adaptable classification categories

Future Development

othman-moumni-abdou has indicated this will be covered in an upcoming article in Alan's document processing series, detailing specific implementation strategies and lessons learned.

See also

Cross-Industry Document Processing

page dédiée →

Validation that production LLM document processing techniques and challenges generalize across industries, with patterns identified in healthcare applying equally to financial, legal, and other structured document domains. Key evidence comes from holofin's confirmation of alan-health's insights.

Universal Production Patterns

Common Challenges Across Industries:

  1. Few-Shot Contamination: External knowledge bleeding into examples affects healthcare, financial, and legal document processing equally
  2. Classification Bottlenecks: Single points of failure in document categorization impact all structured document workflows
  3. Document Quality Issues: Handwriting, poor scans, and image degradation affect all industries processing physical documents
  4. Layout-Based Similarity: Visual structure matching outperforms semantic content matching across domains

Cross-Industry Validation Examples

Healthcare → Financial: edouard-foussier at holofin confirmed that layout-based similarity matching, originally developed for French healthcare documents at alan-health, applies equally to financial document processing.

Technical Principles: Multimodal approaches, pure parsing separation, and evaluation frameworks show consistent value across healthcare and financial document automation.

Why Techniques Generalize

Structural Similarity: Business documents across industries follow similar patterns:

  • Form-based layouts with predictable field arrangements
  • Table structures requiring spatial understanding
  • Template-based designs with visual hierarchy
  • Mixed content types (text, numbers, signatures, stamps)

LLM Behavior Consistency: Core challenges with LLM document processing stem from model characteristics rather than domain specifics:

  • Hallucination tendencies when examples contain external knowledge
  • Superior performance with combined text+image inputs
  • Need for systematic evaluation and quality control

Implementation Benefits

Shared Best Practices: Organizations can adopt proven techniques from other industries rather than reinventing solutions.

Accelerated Development: Cross-industry validation reduces experimentation time and risk when building new document processing systems.

Universal Architecture: Core pipeline components (classification, extraction, validation, evaluation) apply broadly with domain-specific customization only in business logic layers.

Industry-Specific Adaptations

While core techniques generalize, each industry requires customization in:

  • Schema Design: Field types and validation rules specific to document types
  • Regulatory Compliance: GDPR for healthcare, SOX for financial, industry-specific requirements
  • Domain Knowledge: Post-processing enrichment using industry-specific databases and rules
  • Quality Standards: Criticality weights and accuracy thresholds based on business impact

See also

Cross-Industry Validation

page dédiée →

The process of confirming that production AI techniques and insights developed in one industry apply effectively to other sectors, establishing general principles rather than domain-specific solutions. Critical for determining the broader applicability and transferability of production AI practices.

Key Validation Example

Healthcare to Financial Document Processing

othman-moumni-abdou's production insights from alan-health processing French healthcare documents received validation from edouard-foussier at holofin working on financial documents. This cross-industry confirmation established several techniques as broadly applicable:

  • Layout-Based Similarity: Visual structure matching for few-shot examples works across both healthcare and financial documents
  • Multimodal Processing: OCR + image input advantages apply beyond healthcare
  • Few-Shot Contamination Challenges: Human-enriched examples create similar problems across industries
  • Production Evaluation Needs: Systematic backtesting requirements generalize

Significance for Production AI

Technique Generalizability

Cross-industry validation increases confidence that specific production techniques represent fundamental principles rather than domain-specific optimizations. This allows:

  • Faster adoption across industries
  • Higher confidence in implementation decisions
  • Reduced need for domain-specific research
  • Better resource allocation for technique development

Pattern Recognition

Validation across industries helps identify which challenges are universal vs domain-specific:

  • Universal: Document quality issues, classification bottlenecks, evaluation framework needs
  • Domain-Specific: Vocabulary overlap patterns, regulatory requirements, specific document types

Implementation Value

Risk Reduction

Cross-industry validation provides evidence that production techniques will work in new contexts, reducing implementation risk and development time.

Knowledge Transfer

Successful patterns from one industry can be rapidly adapted to others, accelerating production AI development across sectors.

Future Applications

This validation pattern suggests other production AI insights from healthcare, finance, legal, and other document-intensive industries may have broader applicability than initially recognized.

See also

Curated Reference Datasets

page dédiée →

Small, hand-picked collections of high-quality document examples with verified extractions used to bootstrap new document processing categories. Key insight from alan-health: a small, carefully curated reference dataset enables rapid deployment of new document types without requiring large validated pools.

Core Principle

Quality over Quantity: A small set of hand-picked reference documents with verified extractions gets you "surprisingly far" when bootstrapping new categories. Focus on representative, clean examples rather than large volumes of potentially inconsistent data.

Implementation Approach

Manual Curation: Human experts select representative examples covering typical document variations within a category.

Verified Ground Truth: Each reference document has manually validated, high-quality extraction serving as ground truth for evaluation and few-shot selection.

Category Bootstrapping: New document types can be launched with minimal initial data, then improved iteratively as more validated examples accumulate.

Production Benefits

  • Enables rapid deployment of new document processing categories
  • Provides reliable foundation for few-shot-learning without large data requirements
  • Supports evaluation-framework-design with trusted ground truth references
  • Scales naturally as system processes more documents over time

See also

Document Processing Pipeline

page dédiée →

Production-grade system for extracting structured data from documents using LLMs, particularly effective for complex healthcare and insurance documents. Combines OCR transcription with document images for optimal accuracy, as demonstrated by alan-health's 70% automation rate on French healthcare documents.

Architecture Evolution

Text-Only Processing (Initial Approach)

  • Input: OCR Markdown transcription only
  • Output: Structured data extraction
  • Performance: Good baseline accuracy for most documents

Image-Only Processing (Experimental)

  • Input: Document images only, no OCR
  • Findings: Lower accuracy than text-based extraction
  • Problems: Hallucinations where LLM "read" information not actually on documents
  • Conclusion: OCR text provides more reliable parsing foundation than raw pixel interpretation

Multimodal Processing (Current Best Practice)

  • Input: OCR Markdown transcription + document images combined
  • Performance: Outperforms either input alone
  • Benefits:
    • OCR text provides reliable, precise content parsing
    • Images provide visual layout context, field positioning, table structure
    • Combined approach leverages strengths of both modalities

Production Components

Validation Layer

  • Pydantic schema validation for structured output
  • Failed validation triggers human review with structured error messages
  • Best-effort extraction preserved as starting point for reviewers

Human-in-the-Loop Integration

  • Configurable human review for high-stakes document categories
  • Systematic review for new document categories during rollout
  • Online evaluation comparison with manual parsing results

Reference Dataset Strategy

  • Curated examples bootstrap new document categories
  • Small, hand-picked datasets achieve surprising effectiveness
  • Validated documents become potential few-shot examples over time

Future Evolution

Current multimodal capabilities are maturing rapidly. As models improve image-only processing accuracy, the OCR transcription step may become unnecessary, simplifying the pipeline architecture.

See also

Few-Shot Contamination

page dédiée →

Problem in LLM few-shot learning where examples contain information not visible in the source document, teaching the model to hallucinate missing data. Critical production issue identified at alan-health where human-enriched examples corrupt pure document extraction, leading to fabricated values that appear plausible but aren't document-grounded.

Problem Description

Root Cause: Human operators processing documents don't limit themselves to visible content. They:

  • Check other documents in same claim for context
  • Look up healthcare procedure codes and standard prices online
  • Apply domain knowledge not present in document text
  • Make inferences from external systems or databases

Contamination Mechanism: When these "enriched" extractions become few-shot examples, LLMs learn to mimic the enrichment behavior, hallucinating values based on learned patterns rather than document content.

Manifestation: LLM produces plausible-looking but fabricated values for fields where information isn't actually present in the document, because training examples taught it such values "should be there."

Production Impact

Trust Erosion: Users receive extracted data containing hallucinated values that appear legitimate, undermining confidence in automated processing.

Downstream Errors: Hallucinated financial amounts, dates, or codes propagate through business systems causing operational issues.

Detection Difficulty: Hallucinated values often pass schema validation and appear reasonable, making contamination hard to catch without ground truth comparison.

Solution: Architectural Separation

Pure Parsing Phase: Extract only information explicitly visible in document content. Few-shot examples limited to document-grounded extractions only.

Post-Processing Phase: Separate step for enrichment using:

  • Business logic and rules
  • Cross-document context lookup
  • External API calls and database queries
  • Domain knowledge application

Benefits: LLM learns clean document extraction patterns while business enrichment happens in controlled, auditable post-processing step where hallucination risk is eliminated.

Implementation at Scale

alan-health implementing this separation across millions of French healthcare documents to prevent few-shot contamination while maintaining enrichment capabilities through architectural design rather than prompt engineering.

See also

HNSW Indexing

page dédiée →

Hierarchical Navigable Small World (HNSW) indexing technique used for approximate nearest neighbor search in document processing pipelines, trading small precision loss for significant speed improvements when searching large document example pools.

Problem Context

For high-volume document categories with millions of reference examples, exact L2 distance search becomes computationally expensive. alan-health implemented HNSW indexing to maintain performance as their reference document pools grew to massive scale.

Technical Approach

  • Exact L2 Distance: Accurate but slower as example pools grow
  • HNSW Approximate: Small precision trade-off for major speed improvements
  • Use Case: High-volume document categories requiring fast similarity matching

Implementation Benefits

  • Maintains acceptable accuracy for few-shot example selection
  • Scales to millions of reference documents
  • Enables real-time similarity matching in production
  • Supports efficient nearest neighbor queries for document matching

Production Context

Part of alan-health's evolved document processing pipeline, used alongside:

Performance Characteristics

HNSW provides a configurable trade-off between:

  • Search speed (higher with HNSW)
  • Search precision (slightly lower with HNSW)
  • Memory efficiency
  • Scalability to large document pools

See also

Layout-Based Similarity

page dédiée →

Document processing technique that matches few-shot examples based on visual layout structure rather than semantic text content, improving extraction accuracy by providing structurally similar reference documents. Key insight from edouard-foussier at holofin, validated across both financial and healthcare document processing.

Core Principle

Instead of selecting few-shot examples based on textual similarity or semantic content overlap, this approach prioritizes documents with similar visual structures:

  • Table layouts and positioning
  • Field arrangement and spacing
  • Header/footer structures
  • Visual organization patterns

Cross-Industry Validation

Originally identified as effective for financial documents at holofin, this technique validates insights from alan-health's healthcare document processing pipeline. The consistency across industries suggests this is a fundamental principle for multimodal document processing rather than domain-specific optimization.

Implementation Benefits

Improved Extraction Accuracy

Visual layout similarity provides better context for LLM extraction because:

  • Document structure strongly correlates with field positioning
  • Similar layouts indicate similar extraction patterns
  • Visual cues guide field identification more effectively than text similarity

Reduced Few-Shot Contamination

By focusing on layout rather than content, this approach helps avoid semantic bias that can lead to hallucinated values, addressing the few-shot contamination problem identified in production systems.

Technical Implementation

Requires capability to:

  • Analyze document visual structure
  • Create layout-based similarity metrics
  • Select examples based on structural matching rather than content matching
  • Potentially use computer vision techniques for layout analysis

Production Impact

edouard-foussier's confirmation at holofin validates this as a production-ready technique that improves document processing accuracy across different document types and industries, establishing it as a best practice for multimodal document processing systems.

See also

Multimodal Document Processing

page dédiée →

Production approach combining both OCR text transcription and document images as input to LLM systems, outperforming either single-modal approach. Represents evolved best practice as of 2025-2026, demonstrated at scale by alan-health processing French healthcare documents.

Evolution Path

Text-Only → Image-Only → Multimodal

  1. Text-Only Processing: Initial approach using only OCR Markdown transcription as LLM input. Worked well for most documents but lacked visual context.

  2. Image-Only Experiment: Attempted processing with document images alone, skipping transcription. Result: Lower accuracy than text-based extraction, with increased hallucinations where LLMs "read" information not actually present in ambiguous or hard-to-parse images.

  3. Multimodal Combination: Current best practice combining both OCR transcription and document images. Outperforms either input alone.

Why Multimodal Works

Complementary Information Sources:

  • OCR Transcription Provides: Reliable text content that LLMs can parse precisely, reducing ambiguity in character recognition
  • Document Images Provide: Visual layout context, spatial relationships, table structures, presence of stamps/signatures, form organization

Reduced Hallucination Risk: OCR text anchors the extraction to actual document content, while images provide visual verification and context that prevents misinterpretation.

Technical Implementation

Input Format: Send both OCR Markdown transcription AND document image to multimodal LLM models simultaneously.

Processing Strategy: LLM uses text for precise content extraction while referencing image for layout understanding and visual verification.

Production Benefits:

  • Higher extraction accuracy than single-modal approaches
  • Better handling of complex table structures
  • Improved recognition of visual elements (stamps, signatures)
  • Reduced hallucination on ambiguous content

Future Evolution

As of early 2025-2026, multimodal LLMs still require OCR transcription support for optimal accuracy. However, visual capabilities are maturing rapidly - at some point, models may achieve sufficient accuracy with image-only input, potentially eliminating the OCR transcription step.

Production Validation

alan-health's production system processes millions of French healthcare documents using this multimodal approach, achieving 70% automation rates with higher accuracy than previous single-modal systems.

See also

Multimodal LLM Processing

page dédiée →

Advanced document processing approach combining both OCR-extracted text and document images as input to LLMs, achieving superior extraction accuracy compared to either modality alone. Key insight from alan-health: text provides reliable content parsing while images provide crucial visual layout context.

Processing Evolution

Single Modality Limitations

Text-Only Processing: Initial approach using only OCR Markdown transcription worked well for most documents but missed visual context cues.

Image-Only Processing: Experimental approach achieved lower accuracy than text-based extraction, with significant hallucination problems where LLMs would "read" information not actually present in ambiguous or hard-to-parse images.

Optimal Multimodal Approach

Combined OCR + Image: Current best practice providing superior results through:

  • OCR Text: Reliable content that LLMs can parse precisely
  • Document Image: Visual layout context showing field positioning, table structures, stamps, signatures

Technical Advantages

Complementary Information

  • Text provides precise character-level content extraction
  • Images provide spatial relationships and visual formatting context
  • Combined input reduces ambiguity in document structure interpretation

Hallucination Reduction

Image-only processing produced hallucinations where LLMs invented values from ambiguous visual content. OCR text anchors the LLM to actual document content while images provide confirmatory visual context.

Layout Context

Visual elements crucial for accurate extraction:

  • Field positioning relative to labels
  • Table structure and column alignment
  • Presence of signatures, stamps, or annotations
  • Document formatting and visual hierarchy

Future Evolution

Current multimodal capabilities are maturing rapidly. othman-moumni-abdou predicts that eventually models will be accurate enough with just image input, potentially making the OCR transcription step unnecessary as pure visual processing capabilities improve.

Production Implementation

Requirements

  • Multimodal LLM capability (text + image input)
  • OCR pipeline for text extraction
  • Image preprocessing and formatting
  • Validation framework for both modalities

Performance Monitoring

Production systems need evaluation frameworks that can assess:

  • Text extraction accuracy vs image interpretation accuracy
  • Multimodal combination effectiveness
  • Regression detection when updating either OCR or vision components

See also

Production AI Hiring

page dédiée →

Focus area for companies building production LLM systems, requiring engineers who can handle the unique challenges of making AI work reliably on messy real-world data at scale. alan-health actively recruiting for engineers interested in document processing, LLM reliability, and production AI systems.

Required Skills

  • Document processing pipeline development
  • LLM reliability engineering
  • Real-world data handling at scale
  • Production system architecture
  • Evaluation framework design

Challenge Areas

  • Making AI work on messy, real-world data
  • Scaling LLM systems to high-volume production
  • Building reliable evaluation and monitoring systems
  • Handling edge cases and data quality issues
  • Designing human-in-the-loop workflows

Market Demand

The production AI engineering field requires specialized expertise that combines traditional software engineering with deep understanding of LLM behavior, evaluation methodologies, and the unique challenges of deploying AI systems in enterprise environments.

See also

Pure Parsing Separation

page dédiée →

Architectural principle separating pure document extraction (what's visible on the document) from post-processing enrichment (external context, business logic, cross-document inference). Key solution for preventing few-shot-contamination in production document processing systems.

Problem Addressed

When human operators create training examples, they often include information not visible in the source document:

  • Cross-document context from related claims
  • External knowledge (procedure codes, standard prices)
  • Domain expertise and business logic application

Using these "enriched" examples as few-shot training data teaches LLMs to hallucinate missing values, mimicking human inference patterns inappropriately.

Architectural Solution

Two-Stage Processing

  1. Pure Parsing Stage: LLM extracts only what's directly visible in the document
  2. Post-Processing Stage: Separate system applies business logic, external lookups, and cross-document context

Benefits

  • Eliminates Few-Shot Contamination: Training examples contain only document-grounded information
  • Reduces Hallucinations: LLM cannot invent values based on external context
  • Improves Accuracy: Document extraction becomes more reliable and predictable
  • Enables Validation: Pure parsing results can be verified against source documents

Implementation Requirements

Clean Training Data

  • Review existing few-shot examples for external contamination
  • Separate document-visible information from inferred information
  • Create pure parsing examples that contain only document content

Architectural Separation

  • Independent pure parsing component focused solely on document content
  • Separate post-processing pipeline for enrichment and business logic
  • Clear interface between parsing and enrichment stages

Validation Framework

  • Verify pure parsing outputs against source documents
  • Monitor for hallucination patterns indicating contamination
  • Test enrichment logic independently of parsing accuracy

Production Benefits

alan-health identified this separation as crucial for production reliability:

  • More predictable LLM behavior in document extraction
  • Easier debugging when issues arise (parsing vs enrichment problems)
  • Better evaluation capability for each pipeline stage
  • Reduced maintenance overhead from hallucination issues

See also

Pydantic Validation

page dédiée →

Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.

Core Concept

Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.

Basic Implementation Pattern

from pydantic import BaseModel
from typing import Literal

class LLMResponse(BaseModel):
    content: str
    confidence: float
    action_type: Literal["question", "answer", "guess"]
    
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)

Production Applications

Document Processing Pipeline (alan-health)

  • Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
  • Human Review Integration: Structured rejection flows when validation fails
  • Audit Trail: Maintains validation history for compliance requirements
  • Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review

Game Logic (lesphinx)

  • Deterministic Behavior: SphinxAction model ensures AI responses contain required fields
  • Action Classification: action_type field enables branching logic (question vs. guess)
  • Confidence Tracking: Validated confidence scores for game flow control
  • Multilingual Support: Schema validation across language boundaries

Key Validation Patterns

Field-Level Constraints

class DocumentExtraction(BaseModel):
    medical_code: str = Field(min_length=3, max_length=10)
    confidence: float = Field(ge=0.0, le=1.0)  # Between 0 and 1
    review_required: bool

Enum Validation for Controlled Vocabularies

class GameAction(BaseModel):
    action_type: Literal["question", "guess", "end"]
    language: Literal["fr", "en"]

Custom Validators

from pydantic import validator

class MedicalDocument(BaseModel):
    @validator('medical_code')
    def validate_code_format(cls, v):
        if not v.startswith('ICD'):
            raise ValueError('Medical code must start with ICD')
        return v

Error Handling Strategies

Validation Failure Recovery

  1. Immediate Retry: Re-prompt LLM with validation error context
  2. Fallback Schemas: Simpler models when complex validation fails
  3. Human Handoff: Queue for manual review when automated validation consistently fails
  4. Default Values: Safe defaults for non-critical fields

JSON Security

Critical for preventing injection attacks in LLM-generated content:

# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'

# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})

Production Benefits

Reliability

  • Type Safety: Compile-time checking prevents runtime type errors
  • Data Integrity: Ensures downstream systems receive expected data formats
  • Graceful Degradation: Structured error handling when LLM outputs are malformed

Maintainability

  • Schema Evolution: Versioned models enable backward-compatible API changes
  • Documentation: Pydantic models serve as living documentation of data contracts
  • Testing: Easy mock generation for unit tests using model factories

Observability

  • Validation Metrics: Track validation success rates across different LLM providers
  • Error Classification: Categorize validation failures for model improvement
  • Performance Monitoring: Measure validation overhead in processing pipelines

Integration with LLM APIs

Structured Generation

Many LLM providers support JSON mode or schema-guided generation:

# OpenAI JSON mode
response = openai.chat.completions.create(
    model="gpt-4",
    response_format={"type": "json_object"},
    messages=[...]
)

# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)

Response Parsing Pipeline

def process_llm_response(raw_response: str, schema: Type[BaseModel]):
    try:
        # Primary validation attempt
        return schema.model_validate_json(raw_response)
    except ValidationError as e:
        # Log validation error details
        logger.error(f"Validation failed: {e}")
        # Attempt retry or fallback
        return handle_validation_failure(raw_response, schema, e)

Anti-Patterns

Over-Validation

  • Excessive Constraints: Too strict validation can cause high failure rates
  • Rigid Schemas: Inflexible models that break with reasonable LLM variations

Under-Validation

  • Missing Required Fields: Optional fields that should be required for business logic
  • Weak Type Constraints: Using Any or str when specific types are needed

Error Handling Gaps

  • Silent Failures: Catching validation errors without proper logging or recovery
  • Blocking Operations: Synchronous validation that blocks processing pipelines

See also

RAG vs Wiki Pattern

page dédiée →

Fundamental comparison between traditional Retrieval-Augmented Generation (RAG) systems and andrej-karpathy's llm-wiki-pattern, highlighting different approaches to knowledge management and query answering.

Traditional RAG Approach

Process Flow

  1. Upload collection of files
  2. Chunk and embed documents
  3. At query time: retrieve relevant chunks
  4. Generate answer from retrieved fragments
  5. Repeat process for each new query

Characteristics

  • Stateless: Each query starts fresh
  • Rediscovery: Knowledge reconstructed every time
  • Fragment-based: Works with document chunks, not integrated knowledge
  • No accumulation: Understanding doesn't build over time
  • Limited synthesis: Difficult to connect insights across multiple documents

Examples

  • NotebookLM
  • ChatGPT file uploads
  • Most commercial RAG systems
  • Document Q&A tools

LLM Wiki Pattern Approach

Process Flow

  1. Ingest sources into persistent wiki structure
  2. LLM builds and maintains cross-referenced knowledge base
  3. At query time: synthesize from existing wiki pages
  4. File valuable answers back as new wiki content
  5. Knowledge compounds with each interaction

Characteristics

  • Stateful: Persistent knowledge accumulation
  • Pre-synthesized: Understanding built incrementally over time
  • Integration-based: Works with structured, cross-referenced content
  • Compounding: Each addition makes the whole more valuable
  • Rich synthesis: Connections already established across sources

Key Differences

Aspect Traditional RAG Wiki Pattern
Knowledge State Stateless retrieval Persistent accumulation
Processing Time Query-time discovery Ingest-time integration
Cross-reference Ad-hoc during query Pre-established and maintained
Contradictions Discovered per query Flagged and tracked systematically
Synthesis Quality Limited by retrieval Rich, pre-built connections
Maintenance None (static chunks) Automated wiki maintenance

When to Use Each

RAG Works Well For

  • Simple document Q&A
  • One-off queries against large document sets
  • When you don't need knowledge to compound
  • Rapid deployment without setup overhead
  • Documents that rarely need cross-referencing

Wiki Pattern Works Well For

  • Long-term knowledge building
  • Complex synthesis across multiple sources
  • Research that builds over time
  • When contradictions and evolution matter
  • Domains where connections between concepts are valuable

Hybrid Approaches

Some systems might combine both patterns:

  • Wiki pattern for core, frequently-accessed knowledge
  • RAG for supplementary document collections
  • Different layers for different types of content

Performance Implications

RAG

  • Pros: Simple setup, works immediately, scales to large document sets
  • Cons: Repeated processing overhead, limited synthesis depth, no knowledge accumulation

Wiki Pattern

  • Pros: Rich synthesis, compounding value, deep cross-references, maintained consistency
  • Cons: Higher setup cost, requires ongoing LLM maintenance, more complex architecture

See also