Fault-Tolerant AI Systems
Design principles and patterns for building AI systems that continue to function despite component failures, network issues, or degraded performance from external services. Essential for production AI systems where availability is critical.
Core Principles
Graceful Degradation
Systems should provide partial functionality when components fail rather than complete service outage:
- Partial Results: Return available data even if some sources are unavailable
- Feature Fallbacks: Degrade to simpler algorithms when sophisticated ones fail
- Mode-Specific Operation: Allow subsystems to function independently
Configuration Management
Unified configuration prevents drift between ingestion and retrieval phases:
- Centralized Settings: Single source of truth for collection names, endpoints, credentials
- Environment Consistency: Ensure
.env.examplematches actual code expectations - Runtime Validation: Fail fast with clear error messages on misconfiguration
Service Independence
Design components to function with partial service availability:
- Optional Dependencies: Mark non-critical services as optional
- Circuit Breaker Pattern: Detect and route around failing services
- Caching Strategies: Serve stale data when upstream services are unavailable
Real-World Example: Hybrid Retrieval Systems
The assistant-rh project illustrates common fault-tolerance anti-patterns:
Problem
System required both meilisearch and qdrant for lexical searches, even though Meilisearch alone could provide results. When Qdrant failed, users got empty results instead of the lexical matches that were successfully retrieved.
Solution Pattern
def fault_tolerant_search(query, mode="hybrid"):
lexical_results = search_meilisearch(query)
if mode == "lexical":
return lexical_results # Skip vector search entirely
try:
vector_results = search_qdrant(query)
return merge_results(lexical_results, vector_results)
except QdrantException:
logger.warning("Vector search unavailable, falling back to lexical")
return lexical_results
Implementation Strategies
Monitoring and Observability
- Health Checks: Regular probes for dependent services
- Error Metrics: Track failure rates and response times
- Alerting: Notify operators of degraded service modes
Testing for Failure Modes
- Chaos Engineering: Deliberately introduce failures during testing
- Service Mocking: Test behavior when dependencies are unavailable
- Regression Suites: Verify fault tolerance doesn't regress
Configuration Anti-Patterns
Common configuration issues that break fault tolerance:
Hard-Coded Dependencies
# Bad: Hard-coded collection names
def search():
return qdrant_client.search("legi", query)
# Good: Configurable collection names
def search():
collection = config.get("QDRANT_COLLECTION")
return qdrant_client.search(collection, query)
Environment Variable Drift
- Keep
.env.examplesynchronized with actual code expectations - Validate required variables at startup
- Use structured configuration objects rather than direct environment access
See also
- hybrid-retrieval-systems
- system-resilience
- monitoring
- configuration-management
- assistant-rh