~/wiki

flash moe

---
title: Flash-MoE
category: concepts
created: 2026-04-14
updated: 2026-12-22
tags: [flash-moe, mixture-of-experts, metal-programming, quantization, inference-optimization, macbook-inference, c-programming, edge-ai, performance-optimization, fma-kernels, ssd-streaming, tool-calling-reliability, production-quality, performance-benchmarks, qwen3.5-397b, qwen3.5-397b-a17b, json-output-quality, quality-cliff, pure-c-metal, ai-human-collaboration, 24-hour-development, custom-pipeline, objective-c, metal-shaders, 90-experiments, 58-experiments, server-rack-alternative, m3-max, 48gb-ram, 4.36-tokens-per-second, 209gb-disk, no-frameworks, hand-tuned-shaders, fma-kernel-optimization, baseline-comparison, breakthrough-edge-ai, laptop-deployment, github-awesome, technical-paper, performance-table, 2-bit-vs-4-bit, 397b-parameters]
sources: [raw/screenshots/6942F948-8C72-4F62-A8F5-7E73005FB64B_1_105_c.jpeg]
confidence: high
---

# Flash-MoE

A breakthrough inference system that runs the massive Mixture-of-Experts model [qwen3.5-397b](/concepts/qwen3-5-397b) (Qwen3.5-397B-A17B, 397 billion total / 17 billion active parameters) on a **MacBook Pro M3 Max with 48GB RAM** at 4.4+ tokens/second with production-quality output including tool calling. It demonstrates significant advancement in edge AI deployment through a pure C/Metal implementation with comprehensive performance optimization.

> **Source reconciliation (2026-12-22):** The screenshot source explicitly names the model **Qwen3.5-397B-A17B** and the hardware **M3 Max**. Earlier pages mislabeled these as "Qwen2.5" and "M1 Max"; the Qwen3.5 / M3 Max naming is authoritative.

## Technical Architecture

- **Pure C / Objective-C / Metal** — no Python, no ML frameworks. See [pure-c-metal](/concepts/pure-c-metal) and [metal-programming](/concepts/metal-programming).
- The full **209GB model streams from SSD** through a custom Metal compute pipeline.
- Hand-tuned [fma-kernels](/concepts/fma-kernels) provide the current performance edge.
- Built in a 24-hour AI-human collaboration; documented across 90+ experiments.

## Performance Results

| Configuration | tok/s | Quality | Notes |
|---|---|---|---|
| 4-bit experts, FMA kernel | **4.36** | Excellent | Current best. Full tool calling. 209GB on disk. |
| 4-bit experts, baseline | 3.90 | Excellent | Before FMA kernel optimization. |
| 2-bit experts, trust OS | 5.74 | Good* | 120GB on disk. *Breaks JSON/tool calling. |
| 2-bit peak single | 7.05 | Good* | Warm cache burst. *Not suitable for tool use. |

The scatter plot (58 experiments) tracks tokens/second across Q2/Q4 keep/discard configurations, with running-best lines at 7.05 (Q2) and 4.36 (Q4) tok/s. FMA optimization yields roughly +0.46 tok/s (~12%) over the 4-bit baseline.

The 2-bit configurations illustrate the [quality-cliff](/concepts/quality-cliff): faster but emit malformed JSON (`\name\` instead of `"name"`), breaking [tool-calling-reliability](/concepts/tool-calling-reliability).

## See also

- [qwen3.5-397b](/concepts/qwen3-5-397b)
- [quality-cliff](/concepts/quality-cliff)
- [tool-calling-reliability](/concepts/tool-calling-reliability)
- [fma-kernels](/concepts/fma-kernels)
- [pure-c-metal](/concepts/pure-c-metal)
- [metal-programming](/concepts/metal-programming)