Small Language Models
Mis à jour le 2026-04-14Confiance : high
small-modelsparameter-efficiencymodel-compressionedge-deploymentmobile-ai
Language models with <3B parameters designed to achieve competitive performance while maintaining efficiency for resource-constrained environments. Unlike simply scaled-down versions of larger models, small LMs require specialized architectures and training approaches.
Defining characteristics
Scale and constraints
- Parameter count: Typically under 3 billion parameters
- Memory footprint: Must fit within mobile device memory limits (often <1GB)
- Computational efficiency: Optimized for CPU inference and low-power hardware
- Latency requirements: Sub-100ms response times for interactive applications
Architectural differences
- Custom architectures: Purpose-built designs rather than scaled-down transformers
- Efficient operators: Gated short convolution, grouped query attention, etc.
- Parameter sharing: Techniques like tied embeddings to maximize effective capacity
- Operator selection: Hardware-aware choices (e.g., ShortConv for CPU efficiency)
Training methodology
Pre-training strategies
- Overtraining approach: Training on many more tokens than traditional scaling laws suggest
- Example: LFM2.5-350M trained on 28T tokens (much higher than typical ratios)
- Data quality emphasis: Curated, high-quality datasets more important than quantity
- Task-specific pre-training: Focus on domains relevant to deployment scenarios
Post-training pipeline
- Supervised fine-tuning: More critical than for large models due to limited base capabilities
- Preference alignment: Essential for human-aligned behavior with limited knowledge capacity
- Reinforcement learning: Addresses specific small model issues like doom-looping
- Task specialization: Better to excel at narrow tasks than be mediocre at everything
Unique challenges
Technical issues
- Doom looping: Tendency to generate repetitive patterns during reasoning
- Knowledge limitations: Cannot store extensive world knowledge
- Reasoning constraints: Limited capacity for complex multi-step reasoning
- Context handling: Shorter effective context due to parameter constraints
Solutions and mitigations
- N-gram repetition penalty: RL-based approaches to prevent repetitive generation
- On-policy data generation: Generate training data using the model itself
- Agentic approaches: Train models to use external tools and environments
- Specialized architectures: Custom designs optimized for specific hardware
Applications and deployment
Primary use cases
- Mobile AI assistants: On-device processing for privacy and speed
- IoT applications: Embedded systems with strict resource constraints
- Real-time systems: Applications requiring immediate responses
- Offline scenarios: Environments without reliable internet connectivity
Performance considerations
- Inference speed: Optimized for fast token generation
- Memory efficiency: Minimal RAM usage during inference
- Battery life: Low power consumption on mobile devices
- Thermal management: Avoid overheating in compact devices
Examples and implementations
- liquid-ai's LFM series (350M, 1.2B parameters)
- Gemma 3 270M for mobile deployment
- Qwen3.5-0.8B for multimodal edge applications
- Various distilled models from larger foundation models