Hybrid Local-Cloud LLM Architecture
System design pattern that dynamically switches between local and cloud-based language models based on task requirements, cost considerations, and performance needs. Enables optimal resource utilization while maintaining flexibility and cost control.
Core Architecture
Infrastructure Components
Local Inference Layer:
- DGX Spark x2: High-performance local GPU clusters
- Gemma Models: Open-source models optimized for local deployment
- Mac Mini: Always-on coordination server running litellm-proxy
- Tailscale: Secure network layer for remote access
Cloud API Layer:
- Anthropic Claude: High-capability reasoning and analysis
- OpenAI GPT: Backup/comparison API access
- Cost monitoring: Usage tracking per model and task type
LiteLLM Proxy Configuration
Unified interface abstracting model differences:
# Single API endpoint routing to multiple backends
# Automatic fallback between local and cloud
# Cost tracking and usage optimization
# Model comparison for quality assessment
Task Routing Strategy
Local Model Tasks (Cost-Optimized)
- High-volume processing: Screenshot OCR, audio transcription
- Repetitive operations: Content triage, basic classification
- Privacy-sensitive: Personal project documentation
- Batch processing: Large-scale wiki updates, flashcard generation
Cloud Model Tasks (Quality-Optimized)
- Complex reasoning: Architecture decisions, technical analysis
- Novel content: First-time concept extraction from papers
- Quality-critical: Client-facing output, important synthesis
- Fallback scenarios: When local models are insufficient
Implementation Benefits
Cost Optimization
Measured Performance (Live Implementation):
- 80% of triage operations handled locally (Gemma on DGX)
- 20% complex reasoning routed to Anthropic Claude
- Estimated 60% cost reduction vs. cloud-only approach
- Quality comparison enables data-driven model selection
Development Flexibility
API Switching Capability:
- Zero code changes to switch between models
- A/B testing different models on same tasks
- Gradual migration between providers
- Local development without internet dependency
Performance Characteristics
Local Advantages: Lower latency, unlimited usage, privacy control Cloud Advantages: Higher capability, no infrastructure maintenance, latest models
French Developer Workflow Integration
Mac Mini Coordination Server
- Always-on availability: Wiki agent runs continuously
- LiteLLM proxy hosting: Central model routing
- Tailscale coordination: Secure access from anywhere
- Cost monitoring: Track usage across both local and cloud
Mobile Integration
- iPhone → iCloud → Mac Mini: Screenshot ingestion pipeline
- Telegram bot: Cloud models for real-time interaction
- Local processing: Batch operations during off-hours
- Smart routing: Context-aware model selection
Quality Assessment Framework
Model Comparison Methodology
Parallel Processing: Same content scored by both local and cloud models Quality Metrics: Accuracy, relevance, depth of analysis Learning Loop: Results feed back into routing decisions Continuous Optimization: Routing rules improve with usage data
Performance Tracking
task_routing:
screenshot_triage: local (gemma) # 95% accuracy, 0.1s latency
concept_extraction: cloud (claude) # 99% accuracy, 2s latency
wiki_updates: hybrid # Quality threshold determines routing
This architecture enables ai-engineering-wiki to optimize both cost and quality while maintaining the flexibility to adapt as models and requirements evolve.
See also
- litellm-proxy
- ai-engineering-wiki
- intelligent-content-triage
- cost-optimization-strategies