~/wiki

Gated Short Convolution

Confiance : high
shortconvattention-replacementedge-optimizationliquid-ailfmgated-convolutioncpu-optimization

Novel attention mechanism replacement developed by liquid-ai for their liquid-foundation-models, specifically designed to optimize inference performance on edge devices and CPU-bound environments. ShortConv achieves significant speedups while maintaining model quality.

Architecture

Core Design

B x C → [Linear] → [Linear] → [Conv1D] → Gated Output

Components:

  • Gated structure: Controls information flow like attention gates
  • Short convolutions: Limited kernel size for efficiency
  • Linear projections: Standard feedforward transformations
  • Element-wise gating: Selective activation patterns

Integration Pattern

RMSNorm
3:1 ShortConv/GQA  ← Replaces traditional attention
RMSNorm
Feedforward

Performance Advantages

CPU Optimization

  • 2.5x faster decode vs traditional attention mechanisms
  • Linear complexity vs quadratic attention scaling
  • Reduced memory bandwidth requirements
  • Cache-friendly access patterns

Edge Device Benefits

  • Lower computational overhead for mobile deployment
  • Reduced memory pressure during inference
  • Optimized for ARM and x86 CPU architectures
  • Better thermal efficiency on mobile devices

Comparison with Attention Alternatives

Cost Ratios (M4 Max CPU decode)

  • ShortConv: 1.0x baseline
  • SWA (Gemma3): 1.2x overhead
  • GDN (Qwen3.5): 1.5x overhead
  • GLA: 2.0x overhead
  • GQA: 2.2x overhead

Implementation Details

Gated Short Convolution Block

  • Input processing: Linear transformation of input embeddings
  • Convolution: 1D convolution with optimized kernel size
  • Gating mechanism: Learned gates for selective activation
  • Output projection: Linear transformation to model dimensions

Training Considerations

  • Maintains gradient flow during training
  • Compatible with standard transformer training pipelines
  • Requires minimal architectural changes from attention
  • Supports both pre-training and fine-tuning workflows

Applications

LFM2.5 Series

  • LFM2.5-350M: Primary attention replacement
  • Mobile deployment: Optimized for smartphone inference
  • Real-time applications: Sub-100ms response requirements
  • Edge AI: Resource-constrained environments

Use Cases

  • Conversational AI on mobile devices
  • Real-time text processing applications
  • Edge-deployed language understanding
  • Low-latency completion and generation tasks

Technical Trade-offs

Advantages

  • Significant CPU inference speedup
  • Linear computational scaling
  • Reduced memory requirements
  • Hardware-friendly operations

Considerations

  • Novel architecture requiring specialized implementation
  • Different training dynamics vs attention
  • May require architecture-specific optimization
  • Limited long-range dependency modeling vs full attention

See also