~/wiki

Hybrid Local-Cloud LLM Architecture

Confiance : high
hybrid-architecturelocal-llmcloud-llmlitellm-proxycost-optimizationprivacyperformancegemma-modelsanthropic-claudedgx-sparkmac-mini-serverapi-switchingmodel-comparisonfrench-developer-infrastructure

System design pattern that dynamically switches between local and cloud-based language models based on task requirements, cost considerations, and performance needs. Enables optimal resource utilization while maintaining flexibility and cost control.

Core Architecture

Infrastructure Components

Local Inference Layer:

  • DGX Spark x2: High-performance local GPU clusters
  • Gemma Models: Open-source models optimized for local deployment
  • Mac Mini: Always-on coordination server running litellm-proxy
  • Tailscale: Secure network layer for remote access

Cloud API Layer:

  • Anthropic Claude: High-capability reasoning and analysis
  • OpenAI GPT: Backup/comparison API access
  • Cost monitoring: Usage tracking per model and task type

LiteLLM Proxy Configuration

Unified interface abstracting model differences:

# Single API endpoint routing to multiple backends
# Automatic fallback between local and cloud
# Cost tracking and usage optimization
# Model comparison for quality assessment

Task Routing Strategy

Local Model Tasks (Cost-Optimized)

  • High-volume processing: Screenshot OCR, audio transcription
  • Repetitive operations: Content triage, basic classification
  • Privacy-sensitive: Personal project documentation
  • Batch processing: Large-scale wiki updates, flashcard generation

Cloud Model Tasks (Quality-Optimized)

  • Complex reasoning: Architecture decisions, technical analysis
  • Novel content: First-time concept extraction from papers
  • Quality-critical: Client-facing output, important synthesis
  • Fallback scenarios: When local models are insufficient

Implementation Benefits

Cost Optimization

Measured Performance (Live Implementation):

  • 80% of triage operations handled locally (Gemma on DGX)
  • 20% complex reasoning routed to Anthropic Claude
  • Estimated 60% cost reduction vs. cloud-only approach
  • Quality comparison enables data-driven model selection

Development Flexibility

API Switching Capability:

  • Zero code changes to switch between models
  • A/B testing different models on same tasks
  • Gradual migration between providers
  • Local development without internet dependency

Performance Characteristics

Local Advantages: Lower latency, unlimited usage, privacy control Cloud Advantages: Higher capability, no infrastructure maintenance, latest models

French Developer Workflow Integration

Mac Mini Coordination Server

  • Always-on availability: Wiki agent runs continuously
  • LiteLLM proxy hosting: Central model routing
  • Tailscale coordination: Secure access from anywhere
  • Cost monitoring: Track usage across both local and cloud

Mobile Integration

  • iPhone → iCloud → Mac Mini: Screenshot ingestion pipeline
  • Telegram bot: Cloud models for real-time interaction
  • Local processing: Batch operations during off-hours
  • Smart routing: Context-aware model selection

Quality Assessment Framework

Model Comparison Methodology

Parallel Processing: Same content scored by both local and cloud models Quality Metrics: Accuracy, relevance, depth of analysis Learning Loop: Results feed back into routing decisions Continuous Optimization: Routing rules improve with usage data

Performance Tracking

task_routing:
  screenshot_triage: local (gemma)  # 95% accuracy, 0.1s latency
  concept_extraction: cloud (claude)  # 99% accuracy, 2s latency  
  wiki_updates: hybrid  # Quality threshold determines routing

This architecture enables ai-engineering-wiki to optimize both cost and quality while maintaining the flexibility to adapt as models and requirements evolve.

See also