---
title: FMA Kernels
category: concepts
created: 2026-12-21
updated: 2026-12-22
tags: [fma-kernels, fused-multiply-add, metal-programming, performance-optimization, gpu-kernels, flash-moe, hand-tuned-optimization, apple-silicon, inference-acceleration, numerical-precision, baseline-comparison, 4.36-tokens-per-second, 12-percent-improvement, production-deployment, metal-shaders, qwen3.5-397b-a17b, 3.90-baseline, 0.46-speedup, breakthrough-edge-ai, laptop-deployment, m3-max, 58-experiments, performance-table]
sources: [raw/screenshots/6942F948-8C72-4F62-A8F5-7E73005FB64B_1_105_c.jpeg]
confidence: high
---
# FMA Kernels
Fused Multiply-Add compute kernels that combine multiplication and addition operations into a single hardware instruction, providing both performance and numerical precision benefits. Hand-tuned FMA kernels in [flash-moe](/concepts/flash-moe) demonstrate meaningful performance improvements in AI inference workloads.
## Flash-MoE Implementation
In the [flash-moe](/concepts/flash-moe) deployment of [qwen3.5-397b](/concepts/qwen3-5-397b) on a MacBook Pro M3 Max:
- **4-bit experts, FMA kernel:** 4.36 tok/s (current best, Excellent quality)
- **4-bit experts, baseline (no FMA):** 3.90 tok/s
The FMA kernel contributes roughly **+0.46 tok/s (~12%)** over baseline while preserving Excellent output quality, including full tool calling.
## See also
- [flash-moe](/concepts/flash-moe)
- [metal-programming](/concepts/metal-programming)
- [performance-optimization](/concepts/performance-optimization)