---
title: LFM2.5-350M
category: concepts
created: 2026-12-21
updated: 2027-01-05
tags: [liquid-ai, lfm-series, small-models, edge-ai, shortconv, architecture-design, 350m-parameters, extreme-overtraining]
sources: [raw/articles/Everything I Learned Training Frontier Small Models.md]
confidence: high
---
# LFM2.5-350M
liquid-ai's flagship 350 million parameter language model featuring innovative [shortconv](/concepts/shortconv) attention mechanism designed specifically for edge deployment. Part of the [liquid-foundation-models](/concepts/liquid-foundation-models) series, demonstrating that small models require fundamentally different architectures than scaled-down versions of larger models.
## Architecture
- 16 layers with a 3:1 [shortconv](/concepts/shortconv)/GQA ratio.
- RMSNorm, Feedforward blocks, tied linear output embedding.
- Embedding accounts for ~19% of model params; **effective size ≈ 287M**.
- Gated Short Convolution block (Linear → Conv1D → Linear, gated by C) optimized for CPU decode cost — cheapest operator versus SWA (Gemma3), GDN (Qwen3.5), GLA, and GQA on M4 Max CPU decode.
## Training
- **Pre-trained on 28T tokens** — an extreme-overtraining regime far beyond compute-optimal ratios, justified by [test-time-scaling](/concepts/test-time-scaling) (Roberts et al., "Test-Time Scaling Makes Overtraining Compute-Optimal," arXiv:2604.01411, April 2026). The slides note: "More pre-training works, even at the smallest scale!"
- LFM2.5 training recipe: Pre/Mid-training → Supervised Fine-Tuning (generate thinking traces) → Preference Alignment → Reinforcement Learning.
## Related thinking model
- **LFM2.5-1.2B-Thinking** is a related on-device reasoning model ("On-Device Reasoning Under 1GB," Liquid AI blog, January 2026) used to demonstrate [on-policy-data-generation](/concepts/on-policy-data-generation) for DPO and the n-gram repetition penalty that mitigates the [doom-looping-problem](/concepts/doom-looping-problem).
## Benchmarking
- On-device profiling on galaxy-s24-ultra and ryzen-hx-370.
- CPU inference: Llama.cpp, 4-bit quantization, 2K-token input.
- GPU inference: SGLang, 1024-token input / 256-token output, measured vs. concurrency.
## See also
- [liquid-foundation-models](/concepts/liquid-foundation-models)
- [shortconv](/concepts/shortconv)
- [test-time-scaling](/concepts/test-time-scaling)
- [doom-looping-problem](/concepts/doom-looping-problem)
- maxime-labonne
- liquid-ai