~/wiki

Pydantic Validation

Confiance : high
pydantic-validationschema-validationstructured-datatype-checkingerror-handlingdocument-processingllm-output-validationproduction-systemshuman-review-fallbackjson-securityfield-level-validationgame-state-managementstructured-outputdeterministic-behavior

Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.

Core Concept

Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.

Basic Implementation Pattern

from pydantic import BaseModel
from typing import Literal

class LLMResponse(BaseModel):
    content: str
    confidence: float
    action_type: Literal["question", "answer", "guess"]
    
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)

Production Applications

Document Processing Pipeline (alan-health)

  • Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
  • Human Review Integration: Structured rejection flows when validation fails
  • Audit Trail: Maintains validation history for compliance requirements
  • Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review

Game Logic (lesphinx)

  • Deterministic Behavior: SphinxAction model ensures AI responses contain required fields
  • Action Classification: action_type field enables branching logic (question vs. guess)
  • Confidence Tracking: Validated confidence scores for game flow control
  • Multilingual Support: Schema validation across language boundaries

Key Validation Patterns

Field-Level Constraints

class DocumentExtraction(BaseModel):
    medical_code: str = Field(min_length=3, max_length=10)
    confidence: float = Field(ge=0.0, le=1.0)  # Between 0 and 1
    review_required: bool

Enum Validation for Controlled Vocabularies

class GameAction(BaseModel):
    action_type: Literal["question", "guess", "end"]
    language: Literal["fr", "en"]

Custom Validators

from pydantic import validator

class MedicalDocument(BaseModel):
    @validator('medical_code')
    def validate_code_format(cls, v):
        if not v.startswith('ICD'):
            raise ValueError('Medical code must start with ICD')
        return v

Error Handling Strategies

Validation Failure Recovery

  1. Immediate Retry: Re-prompt LLM with validation error context
  2. Fallback Schemas: Simpler models when complex validation fails
  3. Human Handoff: Queue for manual review when automated validation consistently fails
  4. Default Values: Safe defaults for non-critical fields

JSON Security

Critical for preventing injection attacks in LLM-generated content:

# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'

# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})

Production Benefits

Reliability

  • Type Safety: Compile-time checking prevents runtime type errors
  • Data Integrity: Ensures downstream systems receive expected data formats
  • Graceful Degradation: Structured error handling when LLM outputs are malformed

Maintainability

  • Schema Evolution: Versioned models enable backward-compatible API changes
  • Documentation: Pydantic models serve as living documentation of data contracts
  • Testing: Easy mock generation for unit tests using model factories

Observability

  • Validation Metrics: Track validation success rates across different LLM providers
  • Error Classification: Categorize validation failures for model improvement
  • Performance Monitoring: Measure validation overhead in processing pipelines

Integration with LLM APIs

Structured Generation

Many LLM providers support JSON mode or schema-guided generation:

# OpenAI JSON mode
response = openai.chat.completions.create(
    model="gpt-4",
    response_format={"type": "json_object"},
    messages=[...]
)

# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)

Response Parsing Pipeline

def process_llm_response(raw_response: str, schema: Type[BaseModel]):
    try:
        # Primary validation attempt
        return schema.model_validate_json(raw_response)
    except ValidationError as e:
        # Log validation error details
        logger.error(f"Validation failed: {e}")
        # Attempt retry or fallback
        return handle_validation_failure(raw_response, schema, e)

Anti-Patterns

Over-Validation

  • Excessive Constraints: Too strict validation can cause high failure rates
  • Rigid Schemas: Inflexible models that break with reasonable LLM variations

Under-Validation

  • Missing Required Fields: Optional fields that should be required for business logic
  • Weak Type Constraints: Using Any or str when specific types are needed

Error Handling Gaps

  • Silent Failures: Catching validation errors without proper logging or recovery
  • Blocking Operations: Synchronous validation that blocks processing pipelines

See also