Pydantic Validation
Confiance : high
pydantic-validationschema-validationstructured-datatype-checkingerror-handlingdocument-processingllm-output-validationproduction-systemshuman-review-fallbackjson-securityfield-level-validationgame-state-managementstructured-outputdeterministic-behavior
Schema-driven validation approach for ensuring LLM outputs conform to expected data structures and types. Essential for production document processing systems where structured data reliability is critical, as implemented at alan-health and demonstrated in lesphinx.
Core Concept
Pydantic enables structured output from LLMs by defining Python data classes with type annotations, automatic validation, and JSON serialization. This approach transforms unreliable text generation into reliable structured data suitable for production systems.
Basic Implementation Pattern
from pydantic import BaseModel
from typing import Literal
class LLMResponse(BaseModel):
content: str
confidence: float
action_type: Literal["question", "answer", "guess"]
# LLM outputs JSON that gets parsed and validated
response = LLMResponse.model_validate_json(llm_output)
Production Applications
Document Processing Pipeline (alan-health)
- Medical Document Classification: Validates extracted medical codes, confidence scores, and rejection reasons
- Human Review Integration: Structured rejection flows when validation fails
- Audit Trail: Maintains validation history for compliance requirements
- Error Recovery: Graceful handling of malformed LLM outputs with fallback to human review
Game Logic (lesphinx)
- Deterministic Behavior:
SphinxActionmodel ensures AI responses contain required fields - Action Classification:
action_typefield enables branching logic (question vs. guess) - Confidence Tracking: Validated confidence scores for game flow control
- Multilingual Support: Schema validation across language boundaries
Key Validation Patterns
Field-Level Constraints
class DocumentExtraction(BaseModel):
medical_code: str = Field(min_length=3, max_length=10)
confidence: float = Field(ge=0.0, le=1.0) # Between 0 and 1
review_required: bool
Enum Validation for Controlled Vocabularies
class GameAction(BaseModel):
action_type: Literal["question", "guess", "end"]
language: Literal["fr", "en"]
Custom Validators
from pydantic import validator
class MedicalDocument(BaseModel):
@validator('medical_code')
def validate_code_format(cls, v):
if not v.startswith('ICD'):
raise ValueError('Medical code must start with ICD')
return v
Error Handling Strategies
Validation Failure Recovery
- Immediate Retry: Re-prompt LLM with validation error context
- Fallback Schemas: Simpler models when complex validation fails
- Human Handoff: Queue for manual review when automated validation consistently fails
- Default Values: Safe defaults for non-critical fields
JSON Security
Critical for preventing injection attacks in LLM-generated content:
# DANGEROUS - Direct f-string formatting
message = f'{{"content": "{user_input}"}}'
# SAFE - Proper JSON serialization
import json
message = json.dumps({"content": user_input})
Production Benefits
Reliability
- Type Safety: Compile-time checking prevents runtime type errors
- Data Integrity: Ensures downstream systems receive expected data formats
- Graceful Degradation: Structured error handling when LLM outputs are malformed
Maintainability
- Schema Evolution: Versioned models enable backward-compatible API changes
- Documentation: Pydantic models serve as living documentation of data contracts
- Testing: Easy mock generation for unit tests using model factories
Observability
- Validation Metrics: Track validation success rates across different LLM providers
- Error Classification: Categorize validation failures for model improvement
- Performance Monitoring: Measure validation overhead in processing pipelines
Integration with LLM APIs
Structured Generation
Many LLM providers support JSON mode or schema-guided generation:
# OpenAI JSON mode
response = openai.chat.completions.create(
model="gpt-4",
response_format={"type": "json_object"},
messages=[...]
)
# Validate with Pydantic
validated = GameAction.model_validate_json(response.choices[0].message.content)
Response Parsing Pipeline
def process_llm_response(raw_response: str, schema: Type[BaseModel]):
try:
# Primary validation attempt
return schema.model_validate_json(raw_response)
except ValidationError as e:
# Log validation error details
logger.error(f"Validation failed: {e}")
# Attempt retry or fallback
return handle_validation_failure(raw_response, schema, e)
Anti-Patterns
Over-Validation
- Excessive Constraints: Too strict validation can cause high failure rates
- Rigid Schemas: Inflexible models that break with reasonable LLM variations
Under-Validation
- Missing Required Fields: Optional fields that should be required for business logic
- Weak Type Constraints: Using
Anyorstrwhen specific types are needed
Error Handling Gaps
- Silent Failures: Catching validation errors without proper logging or recovery
- Blocking Operations: Synchronous validation that blocks processing pipelines
See also
- llm-integration-patterns
- structured-output
- json-security
- error-handling
- alan-health
- lesphinx