~/wiki

Small Language Models

Mis à jour le 2026-04-14Confiance : high
small-modelsparameter-efficiencymodel-compressionedge-deploymentmobile-ai

Language models with <3B parameters designed to achieve competitive performance while maintaining efficiency for resource-constrained environments. Unlike simply scaled-down versions of larger models, small LMs require specialized architectures and training approaches.

Defining characteristics

Scale and constraints

  • Parameter count: Typically under 3 billion parameters
  • Memory footprint: Must fit within mobile device memory limits (often <1GB)
  • Computational efficiency: Optimized for CPU inference and low-power hardware
  • Latency requirements: Sub-100ms response times for interactive applications

Architectural differences

  • Custom architectures: Purpose-built designs rather than scaled-down transformers
  • Efficient operators: Gated short convolution, grouped query attention, etc.
  • Parameter sharing: Techniques like tied embeddings to maximize effective capacity
  • Operator selection: Hardware-aware choices (e.g., ShortConv for CPU efficiency)

Training methodology

Pre-training strategies

  • Overtraining approach: Training on many more tokens than traditional scaling laws suggest
  • Example: LFM2.5-350M trained on 28T tokens (much higher than typical ratios)
  • Data quality emphasis: Curated, high-quality datasets more important than quantity
  • Task-specific pre-training: Focus on domains relevant to deployment scenarios

Post-training pipeline

  • Supervised fine-tuning: More critical than for large models due to limited base capabilities
  • Preference alignment: Essential for human-aligned behavior with limited knowledge capacity
  • Reinforcement learning: Addresses specific small model issues like doom-looping
  • Task specialization: Better to excel at narrow tasks than be mediocre at everything

Unique challenges

Technical issues

  • Doom looping: Tendency to generate repetitive patterns during reasoning
  • Knowledge limitations: Cannot store extensive world knowledge
  • Reasoning constraints: Limited capacity for complex multi-step reasoning
  • Context handling: Shorter effective context due to parameter constraints

Solutions and mitigations

  • N-gram repetition penalty: RL-based approaches to prevent repetitive generation
  • On-policy data generation: Generate training data using the model itself
  • Agentic approaches: Train models to use external tools and environments
  • Specialized architectures: Custom designs optimized for specific hardware

Applications and deployment

Primary use cases

  • Mobile AI assistants: On-device processing for privacy and speed
  • IoT applications: Embedded systems with strict resource constraints
  • Real-time systems: Applications requiring immediate responses
  • Offline scenarios: Environments without reliable internet connectivity

Performance considerations

  • Inference speed: Optimized for fast token generation
  • Memory efficiency: Minimal RAM usage during inference
  • Battery life: Low power consumption on mobile devices
  • Thermal management: Avoid overheating in compact devices

Examples and implementations

  • liquid-ai's LFM series (350M, 1.2B parameters)
  • Gemma 3 270M for mobile deployment
  • Qwen3.5-0.8B for multimodal edge applications
  • Various distilled models from larger foundation models

See also