~/wiki

Mixture of Experts

Confiance : high
mixture-of-expertsmoemodel-architecturesparse-modelsparameter-efficiencyroutingexpert-networksflash-moe

A neural network architecture pattern that uses multiple specialized sub-networks (experts) with a gating mechanism that routes inputs to the most appropriate experts. This allows for massive parameter scaling while keeping inference costs manageable through sparse activation.

Core Architecture

MoE models contain many expert networks but only activate a subset during inference, enabling models with hundreds of billions of parameters to run efficiently. The routing mechanism determines which experts process each input token based on learned patterns.

Key Components

  • Expert networks: Specialized sub-models handling different types of inputs
  • Gating mechanism: Router that selects which experts to activate
  • Sparse activation: Only a fraction of total parameters used per forward pass
  • Load balancing: Ensures even distribution of work across experts

Practical Deployment

Recent breakthroughs like flash-moe demonstrate that extremely large MoE models (397B parameters) can run locally on consumer hardware through advanced quantization and hardware optimization techniques. The Qwen2.5-397B-A17B model achieves 4+ tokens/second on a MacBook Pro with proper implementation.

Performance Considerations

  • Memory efficiency: Sparse activation reduces memory bandwidth requirements
  • Quantization compatibility: 4-bit expert quantization maintains quality while reducing storage
  • Streaming potential: Large models can stream from SSD during inference
  • Tool calling support: Production-quality capabilities maintained with proper quantization

Implementation Challenges

  • Load balancing: Preventing expert collapse where some experts are never used
  • Communication overhead: Managing data transfer between experts
  • Quantization artifacts: Lower bit widths can break structured outputs (JSON, tool calls)
  • Hardware optimization: Requiring custom kernels for efficient inference

See also