~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Activation Recomputation

page dédiée →

Memory optimization technique that trades computation for memory by recomputing forward pass activations during the backward pass instead of storing them throughout training. Also known as gradient checkpointing, this technique is fundamental to scaling neural network training to larger models and batch sizes.

Core Concept

Fundamental Trade-off

  • Memory Savings: 50-90% reduction in activation memory usage
  • Computational Cost: 15-20% increase in total computation
  • Net Benefit: Enables training larger models or batch sizes that wouldn't fit in memory otherwise

Why It Works

During standard training:

  1. Forward pass stores all intermediate activations
  2. Backward pass uses stored activations to compute gradients
  3. Peak memory occurs when all activations are stored simultaneously

With activation recomputation:

  1. Forward pass stores only selected checkpoint activations
  2. Backward pass recomputes needed activations from checkpoints
  3. Peak memory reduced to checkpoint storage plus recomputation working memory

Implementation Strategy

Checkpointing Approach

  • Checkpoint Selection: Store activations at strategic layer boundaries
  • Segment Recomputation: Recompute activations within segments during backprop
  • Granularity Control: Balance checkpoint frequency vs. recomputation overhead

Typical Checkpoint Placement

For transformer models:

  • Checkpoint at attention block boundaries
  • Store attention outputs and feed-forward outputs
  • Recompute internal attention and FFN activations as needed

Memory Calculation Example

For Llama 3 8B model:

  • Standard Training: ~61.09 GB activation memory
  • With Recomputation: ~6-30 GB activation memory (depending on checkpoint frequency)
  • Total Savings: 50-90% activation memory reduction

Integration with Training Step Anatomy

Forward Pass Modifications

  • Store only checkpoint activations instead of all intermediate results
  • Continue normal forward computation but discard non-checkpoint activations
  • Mark checkpoint boundaries for backward pass reference

Backward Pass Modifications

  • When gradient computation needs missing activation:
    1. Locate nearest stored checkpoint
    2. Recompute forward pass from checkpoint to needed activation
    3. Use recomputed activation for gradient calculation
    4. Discard recomputed activation after use

Memory Dynamic Changes

Changes the typical training-step-anatomy memory patterns:

  • Forward Pass: Lower peak due to limited activation storage
  • Backward Pass: Micro-spikes during recomputation phases
  • Overall: Significantly reduced memory footprint

Advanced Optimization Techniques

Selective Recomputation

Not all activations need recomputation:

  • Cheap Operations: Always recompute (element-wise operations, layer norms)
  • Expensive Operations: Consider checkpointing (attention, large matrix multiplications)
  • Memory-Heavy: Prioritize for recomputation (large activation tensors)

Overlapping Strategies

  • Computation-Communication Overlap: Recompute activations while communicating gradients
  • Pipeline Integration: Coordinate recomputation with pipeline parallel stages
  • Memory Pool Management: Efficiently manage temporary memory for recomputation

Hardware-Specific Tuning

  • GPU Memory Hierarchy: Utilize L2 cache for frequently recomputed activations
  • Tensor Core Optimization: Ensure recomputed operations use optimal data layouts
  • Mixed Precision: Apply appropriate precision for recomputed vs. stored activations

Production Considerations

Implementation Frameworks

  • PyTorch: Built-in torch.utils.checkpoint functionality
  • Nanotron: Production implementation used at Hugging Face
  • Picotron: Educational reference implementations

Configuration Parameters

  • Checkpoint Frequency: How often to store activations
  • Recomputation Granularity: Size of recomputed segments
  • Memory Budget: Target memory usage vs. compute overhead

Monitoring and Debugging

  • Track recomputation overhead in training metrics
  • Monitor memory usage patterns during recomputation phases
  • Profile backward pass timing to optimize checkpoint placement

Distributed Training Integration

Multi-GPU Coordination

  • Coordinate checkpoint placement across tensor parallel ranks
  • Ensure recomputation doesn't create communication bottlenecks
  • Balance memory savings vs. increased computation across devices

Pipeline Parallelism Interaction

  • Coordinate recomputation with pipeline stage boundaries
  • Optimize bubble time during recomputation phases
  • Balance checkpoint storage across pipeline stages

Expert Parallelism Considerations

  • Apply recomputation selectively to expert vs. shared layers
  • Coordinate expert routing with recomputation scheduling
  • Optimize memory usage across expert parallel groups

Mathematical Foundation

Memory Reduction Formula

For L layers with checkpoint every C layers:

  • Standard Memory: O(L × batch_size × sequence_length × hidden_dim)
  • With Checkpointing: O((L/C + C) × batch_size × sequence_length × hidden_dim)
  • Optimal C: √L for balanced memory-computation trade-off

Computational

bfloat16 Precision

page dédiée →

Brain floating-point 16-bit format designed to accelerate machine learning training while maintaining numerical stability. Provides significant speed and memory improvements over full float32 precision with minimal impact on model quality.

Technical Design

Format Structure:

  • 16 bits total: 1 sign bit, 8 exponent bits, 7 mantissa bits
  • Same exponent range as float32 (better than float16)
  • Reduced mantissa precision compared to float32
  • Direct truncation from float32 (no complex conversion)

Key Advantages:

  • 2x memory reduction compared to float32
  • Faster computation on modern GPUs (nvidia-b300, etc.)
  • Better numerical stability than float16
  • Seamless integration with existing training pipelines

Practical Applications

LLM Training Optimization:

  • Used in gpu-mode-paris-2026 hackathon for 10-minute training constraints
  • Enables larger models or batch sizes within memory limits
  • Combined with gradient-accumulation for memory-efficient training
  • Critical for competitive training scenarios requiring maximum speed

Memory Efficiency:

  • Halves memory footprint for model weights and activations
  • Enables training larger models on fixed hardware
  • Reduces data transfer overhead in distributed-training
  • Particularly effective with modern GPU architectures

Implementation Considerations

Training Pipeline Integration:

  • Automatic mixed precision (AMP) frameworks handle conversions
  • Master weights maintained in float32 for stability
  • Gradients computed and accumulated in reduced precision
  • Loss scaling prevents gradient underflow

Numerical Stability:

  • Generally stable for most deep learning applications
  • May require careful tuning for sensitive operations
  • Loss scaling essential for gradient preservation
  • Model-specific validation recommended

Performance Impact

Speed Improvements:

  • Significant acceleration on tensor processing units
  • Reduced memory bandwidth requirements
  • Faster inter-GPU communication in distributed setups
  • Essential for time-constrained training scenarios

Quality Trade-offs:

  • Minimal impact on final model performance for most applications
  • Occasional need for selective float32 operations
  • Benefits typically outweigh precision costs
  • Critical enabler for large-scale training

See also

Distributed Training

page dédiée →

Training neural networks across multiple GPUs or machines to enable larger models, bigger batch sizes, and faster training. Essential for modern LLM development and production ML systems, representing the foundation for ultra-scale AI development.

Core Training Process

Model Replication: Identical copy of model parameters instantiated on each GPU, ensuring synchronized starting points.

Data Parallelism: Each GPU processes different subset of the training batch, maximizing parallel computation efficiency.

Gradient Synchronization: After backward pass, all GPUs exchange and average their computed gradients before parameter updates.

Parameter Updates: Each GPU applies the averaged gradients to its local model copy, maintaining synchronization across all devices.

Technical Implementation

Distributed Data Parallel (DDP)

PyTorch Implementation: Standard approach using torch.nn.parallel.DistributedDataParallel for automatic gradient synchronization.

Communication Backend: Uses NCCL (NVIDIA Collective Communication Library) for high-performance GPU-to-GPU communication.

Process Groups: Each GPU runs in separate process, coordinating through inter-process communication protocols.

Practical Example (32-GPU Setup)

# Each GPU processes portion of global batch
global_batch_size = 128 * 32  # 4096 total examples
local_batch_size = 128        # 128 examples per GPU

# Training loop
for batch in dataloader:
    loss = model(batch)  # Local computation
    loss.backward()      # Local gradient computation
    # Automatic gradient averaging across all GPUs
    optimizer.step()     # Synchronized parameter update

Performance Characteristics

Linear Scaling: Ideal case achieves N× speedup with N GPUs, though communication overhead creates practical limitations.

Memory Efficiency: Combines with gradient-accumulation to simulate larger effective batch sizes without proportional memory increase.

Precision Optimization: Often uses bfloat16 precision to reduce memory usage and increase training speed without significant accuracy loss.

Competition Context

In gpu-mode-paris-2026 hackathon:

  • Hardware: 32× NVIDIA B300 GPUs in coordinated cluster
  • Time Constraint: 10-minute training window maximizes importance of efficient parallelization
  • Optimization Target: Achieve lowest validation loss through optimal distributed resource utilization

Communication Patterns

AllReduce Operations: Primary communication pattern for gradient averaging across all GPUs simultaneously.

Bandwidth Requirements: High-speed interconnects (InfiniBand, NVLink) critical for minimizing communication overhead.

Synchronization Overhead: Trade-off between communication frequency and gradient staleness affects overall training efficiency.

This represents the backbone technology enabling training of modern large language models, from GPT-3 to contemporary 100B+ parameter models that cannot fit on single machines.

See also

Edge AI Optimization

page dédiée →

Specialized techniques for deploying AI models on resource-constrained edge devices, focusing on memory efficiency, latency optimization, and task-specific performance rather than general capabilities.

Core Constraints

Memory-Bound Operations

  • Models must operate within strict memory limits (<3B parameters)
  • Memory bandwidth more limiting than computational power
  • Parameter efficiency critical for deployment success

Latency Requirements

  • Sub-100ms response times required for user-facing applications
  • Fast prefill more important than decode speed optimization
  • Real-time inference constraints shape architecture decisions

Device-Specific Optimization

  • Mobile processors (Galaxy S24 Ultra, Ryzen HX 370)
  • CPU-optimized inference paths
  • Hardware-specific quantization strategies (4-bit with llama.cpp)

Architecture Strategies

Parameter Distribution

  • Optimize embedding layer size (19% vs traditional 63% allocation)
  • Balance between knowledge storage and computational efficiency
  • Effective model size through strategic parameter allocation

Operator Efficiency

  • Gated Short Convolution blocks show 2.5x better cost ratios
  • Replace attention mechanisms with more efficient alternatives
  • Hardware-specific operator optimization (CPU vs GPU paths)

Model Size Targets

  • <1GB models for on-device reasoning (LFM2.5-1.2B-Thinking)
  • Sub-3B parameter counts for memory-bound constraints
  • Task-specific models over general-purpose alternatives

Inference Optimization

CPU Optimization

  • llama.cpp integration with 4-bit quantization
  • Memory-efficient attention alternatives (ShortConv)
  • Optimized operator cost ratios for CPU decode

GPU Batch Processing

  • SGLang integration for concurrent inference
  • Scaling performance with multiple simultaneous requests
  • Input/output token optimization (1024/256 token targets)

Mobile Deployment

  • On-device profiling and optimization
  • Platform-specific performance tuning
  • Battery and thermal management considerations

Training Considerations

Task-Specific Focus

  • Narrow domain optimization over general capabilities
  • Easy adaptation to new domain-specific data
  • Post-training efficiency for specialized tasks

Memory-Aware Training

  • Architecture choices informed by deployment constraints
  • Parameter allocation strategies during training
  • Inference-first design philosophy

Performance Metrics

Latency Benchmarks

  • Sub-100ms response time requirements
  • Prefill speed optimization priorities
  • Real-time inference capability

Memory Efficiency

  • Model size under deployment constraints
  • Runtime memory usage optimization
  • Quantization impact on accuracy vs efficiency

Throughput Scaling

  • Concurrent request handling
  • Batch processing optimization
  • Resource utilization efficiency

See also

Expert Model Loading

page dédiée →

Memory optimization technique demonstrated in apple-siri-architecture where specialized model components are dynamically loaded from storage into RAM on a per-query basis, enabling large-scale AI capabilities on memory-constrained devices.

Technical Approach

Dynamic Loading System

  • NAND-to-RAM transfer: Experts stored in flash storage, loaded as needed
  • Query-specific activation: Different specialists loaded based on request type
  • Memory footprint optimization: Temporary loading reduces permanent RAM usage
  • Response time trade-off: Storage access latency vs. memory conservation

Architecture Benefits

  • Large model capacity: 20B-parameter model on mobile hardware
  • Memory efficiency: Avoid permanent allocation of all model components
  • Specialization: Different experts for different query types
  • Scalability: Can support more experts than would fit in RAM

Implementation Challenges

Performance Considerations

  • Loading latency: Time required to transfer experts from storage
  • Storage wear: Frequent NAND access may impact device longevity
  • Prediction accuracy: Must correctly anticipate which experts to load
  • Caching strategy: Optimizing which experts remain in memory

Technical Requirements

  • Fast storage: High-speed NAND access for reasonable response times
  • Prediction models: Systems to determine expert requirements from queries
  • Memory management: Efficient allocation and deallocation of expert models
  • Error handling: Graceful fallbacks when expert loading fails

Mobile AI Innovation

Resource Constraint Solutions

Expert loading represents creative adaptation to mobile hardware limitations, enabling sophisticated AI without requiring massive RAM allocation.

Privacy Implications

On-device expert loading supports privacy-preserving AI by avoiding cloud-based processing for sensitive queries.

Broader Applications

Edge Computing

The technique could apply to other resource-constrained environments requiring sophisticated AI capabilities.

Cost Optimization

Cloud deployments might use similar approaches to optimize memory costs in serving infrastructure.

Future Developments

  • Faster storage technologies: Reducing loading latency
  • Better prediction models: More accurate expert selection
  • Hybrid approaches: Combining on-device and cloud experts
  • Cross-platform adaptation: Applying to other mobile and edge platforms

See also

Quantization-Aware Training (QAT) implementation for Gemma 4 models that achieves significant memory reduction while preserving performance. Represents advancement in efficient model deployment for resource-constrained environments.

Performance Characteristics

Memory Reduction: ~4x less memory usage compared to standard Gemma 4 models while maintaining comparable performance.

Mobile Optimization: Gemma 4 E2B variant fits in approximately 1GB using specialized mobile quantization format.

Performance Preservation: QAT training maintains model capabilities despite aggressive quantization.

Technical Implementation

Quantization-Aware Training: Models trained with quantization effects incorporated during training phase, enabling better preservation of capabilities compared to post-training quantization.

Mobile Format: Specialized quantization format optimized for mobile and edge deployment scenarios.

Hardware Integration: Optimized for deployment across various hardware configurations with limited memory.

Integration Support

llama.cpp Compatibility: Gemma 4 MTP merged into llama.cpp for faster decoding when paired with QAT checkpoints.

Ecosystem Support: Broad compatibility with existing inference frameworks and serving infrastructure.

Impact

Enables deployment of advanced language models in previously infeasible environments, advancing democratization of AI capabilities through efficient resource utilization.

See also

Gradient Accumulation

page dédiée →

Memory optimization technique that enables training with larger effective batch sizes without proportionally increasing memory usage. Achieves this by processing data in smaller micro-batches and accumulating their gradients before performing parameter updates.

Core Mechanism

Problem Solved: GPU memory limitations prevent training with optimal batch sizes (e.g., 128 examples requiring 65GB when GPU has only 40GB available).

Solution Approach: Split large batch into smaller micro-batches, process sequentially while accumulating gradients, then perform single parameter update with accumulated gradients.

# Traditional approach (memory overflow)
loss = model(batch_128_examples)  # OOM Error

# Gradient accumulation approach
total_loss = 0
for micro_batch in split_in_4(batch_128_examples):  # 32 examples each
    loss = model(micro_batch)  # 8.75GB per micro-batch
    loss.backward()  # Accumulate gradients
    total_loss += loss
optimizer.step()  # Update with accumulated gradients

Benefits of Larger Effective Batches

Gradient Stability: Averaging over more examples reduces gradient noise, leading to smoother optimization landscapes.

Better Generalization: Models see more data diversity per training step, improving ability to generalize to unseen data.

Faster Convergence: Fewer total optimization steps required to reach target performance due to more informative gradient estimates.

Implementation in Practice

Hyperparameter Configuration: Common configurations use 4-8 gradient accumulation steps, effectively multiplying batch size by that factor.

Memory vs. Compute Trade-off: Exchanges additional forward pass computations for reduced memory requirements, enabling training of larger models on constrained hardware.

Distributed Training Integration: Works synergistically with distributed data parallel (DDP) training, where each GPU accumulates gradients independently before cross-GPU synchronization.

Competition Applications

In the gpu-mode-paris-2026 hackathon context:

  • Standard Configuration: 4 micro-steps accumulating to effective batch size of 128
  • Memory Constraints: Enables training on NVIDIA B300 GPUs without memory overflow
  • Performance Optimization: Critical for achieving competitive validation loss within 10-minute training window

This technique represents a foundational optimization in modern deep learning, used universally in training large language models like GPT, LLaMA, and other transformer architectures.

See also

Inference Optimization

page dédiée →

The field of techniques and strategies to reduce computational cost, memory usage, and latency when running large transformer models in production. Critical for deploying powerful models at scale in real-world applications where cost and performance constraints must be balanced against model capability.

Fundamental Challenges

According to lilian-weng's analysis building on pope-et-al-2022, inference challenges stem from two primary factors beyond just increasing model size:

  1. memory-bandwidth-bottleneck: The rate at which data can be transferred between memory and processing units becomes the limiting factor
  2. autoregressive-generation: Sequential token generation prevents effective parallelization strategies

Core Optimization Strategies

Model Compression

  • quantization: Reducing numerical precision of parameters and activations
  • pruning: Removing less important parameters or connections
  • knowledge-distillation: Training smaller student models to replicate larger teacher behavior

Architecture Optimization

  • attention-optimization: Improving computational and memory efficiency of attention mechanisms
  • Sparse attention patterns: Reducing quadratic scaling of attention computation
  • Key-value caching: Optimizing memory access patterns in autoregressive generation

Hardware Optimization

  • Mixed precision training: Leveraging different numerical precisions for different operations
  • Memory layout optimization: Improving data access patterns
  • Parallel processing strategies: Maximizing utilization of available compute resources

Production Considerations

Real-world deployment requires balancing multiple constraints:

  • Latency requirements: Response time expectations
  • Memory limitations: Available RAM and VRAM constraints
  • Cost optimization: Computational expense vs. model capability
  • Accuracy preservation: Maintaining model performance through optimization

Research Evolution

The field has evolved from simple model size reduction to sophisticated techniques that maintain model capability while dramatically reducing resource requirements. Current research focuses on finding optimal trade-offs between efficiency and performance.

See also

LLM Scaling Techniques

page dédiée →

Methods and strategies for training increasingly large language models, encompassing both model architecture scaling and training infrastructure scaling. Critical for developing state-of-the-art AI systems that require coordination across hundreds to thousands of GPUs.

Batch Size Scaling Evolution

The LLM training community has seen dramatic increases in batch sizes over time, reflecting improved distributed training capabilities:

  • Llama 1: ~4M tokens per batch, 1.4 trillion total training tokens
  • DeepSeek: ~60M tokens per batch, 14 trillion total training tokens
  • DeepSeek-V3/R1: Dynamic scaling from 3,072 to 15,360 input sequences during initial 469B tokens, then maintained at 15,360

Token-Based Measurement

Modern LLM training reports batch sizes in tokens rather than samples to maintain independence from sequence length:

batch_size_tokens = batch_size_samples × sequence_length

This standardization enables consistent comparison across different model architectures and training configurations.

Three-Challenge Framework

All scaling techniques address one or more of three fundamental challenges identified in the ultra-scale-playbook:

1. Memory Usage (Hard Constraint)

  • Nature: If training step doesn't fit in memory, training cannot proceed
  • Components: Model weights, gradients, optimizer states, activations
  • Solutions: memory-optimization, activation-recomputation, gradient accumulation

2. Compute Efficiency

  • Goal: Maximize hardware utilization by reducing idle time
  • Challenges: Data transfer overhead, GPU synchronization delays
  • Solutions: Kernel fusion, mixed precision, optimized data pipelines

3. Communication Overhead

  • Impact: Inter-GPU communication keeps hardware idle
  • Strategy: Optimize intra-node (fast) vs inter-node (slower) bandwidth usage
  • Approach: Overlap communication with computation whenever possible

Five-Dimensional Parallelism

Modern LLM training requires coordination across multiple parallelization dimensions:

  1. Data Parallelism: Distribute training data across GPUs
  2. Tensor Parallelism: Split model tensors across devices
  3. Pipeline Parallelism: Distribute model layers across pipeline stages
  4. Context Parallelism: Handle sequences longer than single-device memory
  5. Expert Parallelism: Distribute mixture-of-experts layers

Trade-off Optimization

Scaling techniques often involve trading one resource against another:

  • Computation vs Memory: activation-recomputation trades 15-20% additional compute for 50-90% memory reduction
  • Communication vs Parallelism: Higher parallelism increases communication overhead
  • Batch Size vs Convergence: Larger batches improve hardware utilization but may affect convergence properties

Empirical Research Foundation

The ultra-scale-playbook provides data-driven optimization through over 4,000 scaling experiments:

  • Up to 512 GPUs tested
  • Systematic measurement of throughput and GPU utilization
  • Normalized performance metrics across different model sizes
  • Reproducible benchmarking methodology

Batch Size Sensitivity and Optimization

Convergence Considerations

Batch size affects model convergence in complex ways:

  • Small batches: Quick initial progress but noisy gradients prevent optimal final performance
  • Large batches: Accurate gradients but potentially slower convergence and compute waste
  • Sweet spot: Typically 4-60 million tokens per batch for recent LLMs

Dynamic Scaling Strategies

Advanced training runs employ dynamic batch size scaling:

  • Start with smaller batches for rapid initial convergence
  • Gradually increase batch size as training progresses
  • Balance convergence speed with computational efficiency

Memory-First Design Philosophy

Modern scaling techniques prioritize memory optimization since memory represents a hard constraint unlike compute or communication which can be optimized without preventing training:

  • Memory prediction tools for capacity planning
  • Systematic memory component analysis
  • training-step-anatomy understanding for optimization
  • Empirical memory profiling for validation

See also

Memory Optimization

page dédiée →

Techniques and strategies for managing GPU memory efficiently during neural network training, particularly for large language models where memory constraints often represent the primary bottleneck in scaling training to larger models and batch sizes.

Memory as Hard Constraint

Memory represents a hard constraint in LLM training - if a single training step doesn't fit in GPU memory, training simply cannot proceed. This makes memory optimization the most critical aspect of distributed training, as established in the ultra-scale-playbook.

Unlike compute or communication which can be optimized for efficiency, memory has absolute limits that must be respected for training to function at all.

Four Memory Components

GPU memory during training consists of four primary components:

1. Model Weights

  • Neural network parameters
  • Stored in various precisions (FP32, BF16, FP8)
  • Size determined by model architecture

2. Gradients

  • Computed during backward pass
  • Typically same size as model weights
  • Temporarily stored before optimization step

3. Optimizer States

  • Often the largest component
  • Adam optimizer stores momentum and variance (2x model size)
  • Can dominate total memory usage

4. Activations

  • Intermediate values from forward pass
  • Needed for gradient computation during backward pass
  • Can be traded for computation through activation-recomputation

Memory Profiling and Analysis

Empirical Measurement Approach

The ultra-scale-playbook emphasizes empirical memory profiling over theoretical calculations:

  • PyTorch Memory Profiler: Step-by-step memory allocation tracking
  • Dynamic Patterns: Memory usage varies significantly during training steps
  • First Step Anomaly: Initial step shows different patterns due to caching allocator preparation

Training Step Memory Anatomy

  1. Forward Pass: Activations build up progressively
  2. Backward Pass: Gradients accumulate while activations are cleared
  3. Optimization: All gradients needed, optimizer states updated

Memory Optimization Techniques

Precision Management

  • Mixed Precision Training: Use BF16/FP16 instead of FP32
  • FP8 Training: Cutting-edge precision for maximum memory savings
  • Gradient Scaling: Maintain numerical stability with lower precision

Activation Management

  • activation-recomputation: Trade computation for memory (50-90% memory reduction)
  • Gradient Checkpointing: Strategic activation storage points
  • Progressive Clearing: Clear activations as soon as gradients computed

Optimizer Optimization

  • zero-optimizer: Partition optimizer states across devices
  • AdamW vs Adam: More memory-efficient optimizer variants
  • State Precision: Lower precision for optimizer states

Batch Size Management

  • gradient-accumulation: Simulate larger batches without memory increase
  • Micro-batching: Process smaller chunks within larger logical batches
  • Dynamic Batching: Adjust batch size based on sequence length

Memory Prediction and Tools

Theoretical Calculation

Memory usage can be estimated from:

  • Tensor shapes (batch size, sequence length, hidden dimensions)
  • Precision formats (4 bytes for FP32, 2 for BF16, 1 for FP8)
  • Model architecture parameters

Empirical Tools

  • Memory Prediction Tools: Hugging Face's memory estimation widgets
  • Profiling Dashboards: Real-time memory usage visualization
  • Benchmarking Suites: Systematic memory usage measurement

Memory Fragmentation and Allocation

PyTorch Caching Allocator

  • Pre-allocates memory blocks to speed up subsequent allocations
  • Causes first step anomaly in memory patterns
  • Can lead to fragmentation reducing usable memory

CUDA Kernel Overhead

  • Kernels typically require 1-2 GB of GPU memory
  • Constant overhead independent of model size
  • Must be factored into memory budget

Trade-offs and Strategies

Computation-Memory Trade-offs

  • Recomputation: Use more computation to reduce memory storage
  • Batching: Larger batches improve efficiency but increase memory
  • Precision: Lower precision saves memory but may affect convergence

Memory-Communication Balance

  • Smaller models per device reduce memory but increase communication
  • Optimal balance depends on interconnect bandwidth
  • Different strategies for intra-node vs inter-node communication

Scaling Implications

Memory optimization becomes increasingly critical at scale:

  • Single GPU: Focus on activation recomputation and mixed precision
  • Multi-GPU: Add optimizer state sharding and gradient compression
  • Ultra-Scale: Combine all techniques with sophisticated parallelism strategies

Understanding memory patterns through empirical profiling enables informed decisions about which optimization techniques to apply for specific training configurations.

See also


Model Compression

page dédiée →

The field of techniques for reducing the size and computational requirements of neural networks while maintaining performance. Critical for deploying large models in resource-constrained environments and reducing inference costs in production systems.

Technical Foundation

As systematically analyzed by lilian-weng, model compression is essential for addressing inference-optimization challenges, particularly the memory-bandwidth-bottleneck and constraints imposed by autoregressive-generation in large transformer models.

Core Compression Techniques

quantization

Reducing numerical precision of model parameters and activations:

  • 8-bit quantization: ~4x size reduction with minimal accuracy loss
  • 4-bit quantization: ~8x size reduction requiring careful implementation
  • Mixed precision: Balancing compression with accuracy preservation
  • Directly addresses memory bandwidth constraints by reducing data transfer requirements

pruning

Removing less important model components:

  • Unstructured pruning: Removing individual parameters based on magnitude or importance
  • Structured pruning: Removing entire neurons, channels, or blocks
  • Sparse models: Maintaining connectivity patterns while reducing active parameters
  • Enables hardware acceleration through specialized sparse computation

knowledge-distillation

Training smaller models to replicate larger model behavior:

  • Teacher-student framework: Large model guides smaller model training
  • Soft targets: Using probability distributions rather than hard classifications
  • Feature matching: Aligning intermediate representations between models
  • Enables deployment of powerful model capabilities in constrained environments

Optimization Objectives

Memory Efficiency

  • Reducing model size for storage and RAM requirements
  • Enabling deployment on edge devices with limited memory
  • Addressing memory bandwidth bottlenecks in inference

Computational Efficiency

  • Reducing FLOPs required for inference
  • Improving throughput and reducing latency
  • Enabling real-time applications with strict timing constraints

Energy Efficiency

  • Reducing power consumption for mobile deployment
  • Extending battery life in portable devices
  • Reducing operational costs in data center deployment

Implementation Strategies

Progressive Compression

  • Gradually applying compression techniques to monitor accuracy impact
  • Starting with less aggressive settings and increasing compression
  • Allows finding optimal points in accuracy-efficiency trade-off space

Multi-technique Combination

  • Applying quantization, pruning, and distillation together
  • Techniques can be complementary when properly orchestrated
  • Requires careful coordination to avoid compounding accuracy losses

Hardware-Aware Compression

  • Tailoring compression to target deployment hardware
  • Leveraging hardware-specific optimizations and constraints
  • Ensuring compressed models can efficiently utilize available resources

Challenges and Trade-offs

Accuracy Preservation

  • Maintaining model performance while reducing complexity
  • Different tasks and architectures have varying compression tolerance
  • Requires careful evaluation and validation processes

Hardware Compatibility

  • Ensuring compressed models work efficiently on target hardware
  • Software framework support for optimized compressed model formats
  • Balancing theoretical compression with practical deployment benefits

Development Complexity

  • Additional engineering effort to implement and validate compression
  • Need for specialized tools and frameworks
  • Increased testing and validation requirements

See also

QAT Quantization

page dédiée →

Quantization-Aware Training (QAT) is an advanced optimization technique that incorporates quantization effects during the training process rather than applying quantization post-hoc. This approach enables dramatic memory reduction while preserving model performance, as demonstrated by gemma-4's achievement of ~4x memory reduction.

Core Methodology

Training Integration: Quantization effects are simulated during training, allowing the model to learn to compensate for precision loss inherent in quantized representations.

Precision Optimization: Balances numerical precision with memory efficiency by training weights to be robust to quantization errors.

Hardware Awareness: Training process considers target deployment hardware constraints, optimizing for specific memory and computational limitations.

Performance Characteristics

Memory Efficiency: Achieves approximately 75% memory reduction compared to full-precision models while maintaining competitive performance metrics.

Quality Preservation: Unlike post-training quantization, QAT maintains model capabilities by learning to work within quantization constraints during training.

Deployment Viability: Enables deployment in severely resource-constrained environments where full-precision models would be infeasible.

Implementation Benefits

Mobile Deployment: Makes advanced AI models viable for smartphone and edge device deployment with ~1GB memory footprints.

Infrastructure Costs: Dramatically reduces serving costs and hardware requirements for model deployment at scale.

Accessibility: Democratizes access to advanced AI capabilities in environments with strict resource constraints.

Technical Challenges

Training Complexity: Requires careful tuning of quantization parameters and training schedules to achieve optimal performance.

Hardware Specificity: Optimal quantization strategies may vary significantly across different target deployment hardware.

Validation Requirements: Extensive testing needed to ensure quantized models maintain acceptable performance across diverse use cases.

Industry Impact

Edge AI Revolution: Enables practical deployment of sophisticated models on consumer devices and IoT hardware.

Cost Optimization: Provides pathway for significant infrastructure cost reduction in large-scale AI deployments.

Innovation Catalyst: Opens new possibilities for AI applications in resource-constrained environments previously considered infeasible.

See also

Quantization

page dédiée →

The process of reducing the numerical precision of neural network parameters and activations from higher precision formats (like FP32 or FP16) to lower precision formats (like INT8 or INT4). A critical technique in model-compression for reducing memory usage, improving inference speed, and enabling deployment on resource-constrained hardware.

Technical Foundation

As analyzed by lilian-weng, quantization is one of the core strategies in inference-optimization for addressing the memory-bandwidth-bottleneck that constrains large transformer model deployment.

Types of Quantization

Post-Training Quantization (PTQ)

  • Applied to already-trained models without additional training
  • Faster to implement but may have larger accuracy drops
  • Suitable for models with sufficient redundancy

Quantization-Aware Training (QAT)

  • Incorporates quantization simulation during training process
  • Better accuracy preservation but requires more computational resources
  • Model learns to be robust to quantization effects

Precision Levels

8-bit (INT8)

  • Reduces model size by ~4x compared to FP32
  • Generally maintains good accuracy with proper calibration
  • Well-supported across hardware platforms

4-bit (INT4)

  • Aggressive compression reducing size by ~8x
  • Requires careful implementation to maintain accuracy
  • Increasingly supported in modern inference frameworks

Mixed Precision

  • Different layers or operations use different precisions
  • Balances compression with accuracy preservation
  • Allows fine-tuning of the precision-accuracy trade-off

Implementation Considerations

Calibration Dataset

  • Representative data used to determine quantization parameters
  • Critical for maintaining model accuracy
  • Should reflect actual deployment data distribution

Quantization Schemes

  • Symmetric: Zero point is at the center of the range
  • Asymmetric: Zero point can be offset for better range utilization
  • Per-channel vs. per-tensor: Granularity of quantization parameters

Memory and Performance Benefits

Memory Reduction

  • Direct reduction in model size proportional to precision decrease
  • Enables deployment on resource-constrained devices
  • Addresses memory-bandwidth-bottleneck by reducing data transfer requirements

Speed Improvements

  • Lower precision arithmetic can be computed faster
  • Hardware-specific optimizations for quantized operations
  • Reduced memory access time due to smaller data sizes

Energy Efficiency

  • Lower precision operations consume less energy
  • Particularly important for edge deployment
  • Extends battery life in mobile applications

Challenges and Limitations

Accuracy Degradation

  • Some accuracy loss is typically unavoidable
  • Certain model architectures more sensitive to quantization
  • Requires careful evaluation of accuracy-efficiency trade-offs

Hardware Support

  • Not all hardware platforms support all quantization schemes
  • Software frameworks may have varying levels of optimization
  • Need to match quantization approach to deployment target

See also

SWA Attention

page dédiée →

Sliding Window Attention mechanism implemented in gemma-3-270m architecture as part of modern efficiency-focused attention innovations. Provides memory-efficient attention computation by limiting the attention window to a fixed size rather than full sequence attention.

Architecture Implementation

Gemma 3 270M Configuration

  • Layer structure: 24 layers total
  • Ratio: 3:1 GDN/Gated Attention with SWA
  • Parameter allocation: 63% of total model parameters
  • Effective size: ~100M parameters despite 270M total

Sliding Window Mechanism

Memory Efficiency

  • Limits attention computation to a fixed window size
  • Reduces memory requirements for long sequences
  • Maintains performance while decreasing computational complexity
  • Enables processing of longer contexts with constrained resources

Computational Benefits

  • Linear memory scaling instead of quadratic with sequence length
  • Improved cache efficiency for edge deployment
  • Reduced latency for real-time applications

Performance Analysis

CPU Inference Cost

In maxime-labonne's M4 Max CPU benchmarking, SWA demonstrates moderate computational overhead compared to shortconv's minimal cost profile but significantly better than traditional full attention mechanisms.

Edge Deployment Advantages

  • Suitable for memory-constrained environments
  • Predictable memory usage regardless of input length
  • Fast prefill performance critical for edge applications

Comparison with Other Mechanisms

Efficiency Ranking (CPU decode cost)

  1. shortconv - Lowest cost, optimized for CPU
  2. SWA - Moderate cost, memory efficient
  3. GDN - Higher cost, multimodal capabilities
  4. GLA/GQA - Variable cost depending on configuration

Applications

Edge Model Deployment

Particularly effective for:

  • Real-time text processing
  • Memory-constrained devices
  • Long document processing with limited resources
  • Applications requiring predictable latency

See also

Ultra-Scale Playbook

page dédiée →

Comprehensive 240+ page methodology and knowledge base for scaling LLM training from single GPUs to thousands of coordinated GPUs, developed by hugging-face through systematic empirical research. Represents the first open-sourcing of previously proprietary distributed training knowledge held within elite industry labs.

Knowledge Democratization Mission

The Ultra-Scale Playbook addresses a critical gap in the AI training ecosystem: while foundation models are openly available, the knowledge and techniques for training them at scale remained "well kept within a handful of big industry labs." This comprehensive resource lifts the veil on distributed training methodologies that were previously proprietary.

Three-Pillar Foundation

The playbook is built on three complementary foundations:

1. Theoretical Understanding

  • Quick introductions to concepts and methods
  • High-level explanations of advantages and limitations
  • Memory breakdown analysis for Transformer models
  • Understanding of when and why memory constraints occur

2. Clear Code Implementations

  • picotron: Educational implementations in single, self-contained files for learning
  • nanotron: Production-ready codebase used at Hugging Face
  • Theory-to-code translations revealing implementation details and edge cases

3. Real Training Efficiency Benchmarks

  • Over 4,000 systematic scaling experiments
  • Up to 512 GPU cluster configurations tested
  • Infrastructure-specific optimization guidance
  • Reproducible performance measurements

Three Core Challenges Framework

All distributed training techniques address one or more of these fundamental challenges:

  1. Memory Usage: Hard constraint - if a training step doesn't fit in memory, training cannot proceed
  2. Compute Efficiency: Maximizing hardware utilization, minimizing idle time and data transfer delays
  3. Communication Overhead: Reducing inter-GPU communication that keeps devices idle, optimizing bandwidth usage

Memory as Hard Constraint

The playbook establishes memory as the primary bottleneck in LLM training, consisting of four critical components:

  • Model Weights: Parameters of the neural network
  • Gradients: Computed during backward pass
  • Optimizer States: Often the largest component (e.g., Adam momentum and variance)
  • Activations: Intermediate values needed for gradient computation

Training Step Anatomy

Detailed analysis of what happens during a single training step:

  1. Forward Pass: Activations build up as inputs pass through layers
  2. Backward Pass: Gradients computed while activations progressively cleared
  3. Optimization Step: All gradients needed, optimizer states updated

First Step Anomaly

The first training step exhibits different memory patterns due to PyTorch caching allocator preparation work. This can lead to successful first steps followed by OOM failures in subsequent steps due to optimizer state buildup.

Batch Size Evolution

Modern LLM training uses token-based batch sizes for sequence-length independence:

  • Llama 1: ~4M tokens per batch, 1.4 trillion total tokens
  • DeepSeek: ~60M tokens per batch, 14 trillion total tokens
  • Sweet Spot: 4-60 million tokens per batch for current LLM training

Educational Philosophy

Combines systematic empirical research with educational accessibility. The approach prioritizes understanding over optimization, making complex distributed training concepts accessible through clear explanations, focused implementations, and real-world benchmarking data.

Empirical Research Foundation

Built on systematic experimentation rather than purely theoretical analysis:

  • Over 4,100 distributed experiments (16k+ including test runs)
  • Systematic scanning of distributed training layouts and model sizes
  • Infrastructure-specific optimization insights
  • Reproducible benchmarking methodologies

Impact

Represents a fundamental shift in knowledge sharing for AI training, moving previously proprietary expertise into the open-source domain. Enables researchers and practitioners to understand and implement ultra-scale training without starting from scratch or reverse-engineering techniques from scattered papers.

See also


ZeRO Optimizer

page dédiée →

Zero Redundancy Optimizer (ZeRO) is an advanced memory optimization technique for distributed training that eliminates memory redundancy by partitioning optimizer states, gradients, and parameters across devices while maintaining training efficiency.

Core Problem: Memory Redundancy

Traditional Data Parallelism Issues

In standard distributed training:

  • Each GPU maintains complete copy of model parameters
  • Each GPU stores full optimizer states (often 2-3x parameter size)
  • Each GPU accumulates complete gradient set
  • Result: Massive memory redundancy across devices

Memory Components

For a model with P parameters using Adam optimizer:

  • Model parameters: P values
  • Gradients: P values
  • Optimizer states: 2P values (momentum + variance)
  • Total per GPU: 4P values × number of GPUs

ZeRO Stages

Stage 1: Optimizer State Partitioning

  • Partition: Optimizer states across devices
  • Memory reduction: 4x reduction for Adam optimizer
  • Communication: Gather required states during optimization
  • Benefit: Significant memory savings with minimal overhead

Stage 2: Gradient Partitioning

  • Partition: Gradients in addition to optimizer states
  • Memory reduction: 8x reduction total
  • Communication: All-reduce only assigned gradient partitions
  • Synchronization: Gradients distributed and synchronized efficiently

Stage 3: Parameter Partitioning

  • Partition: Model parameters across devices
  • Memory reduction: Linear with number of devices
  • Communication: Gather parameters as needed for forward/backward
  • Complexity: Most aggressive but requires careful implementation

Implementation Strategy

Dynamic Parameter Management

Stage 3 requires sophisticated parameter handling:

  1. Forward pass: Gather required parameters just before computation
  2. Computation: Execute with temporarily assembled parameters
  3. Cleanup: Discard non-local parameters to free memory
  4. Backward pass: Repeat gathering for gradient computation

Communication Optimization

  • Overlap: Hide parameter gathering with computation
  • Prefetching: Anticipate parameter needs for next layers
  • Bucketing: Group small parameters for efficient communication

Memory Efficiency Gains

Theoretical Reductions

For N devices:

  • Stage 1: Memory per device = (P + P + 2P/N) = (2P + 2P/N)
  • Stage 2: Memory per device = (P + P/N + 2P/N) = (P + 3P/N)
  • Stage 3: Memory per device = (P/N + P/N + 2P/N) = 4P/N

Practical Benefits

  • Larger models: Train models that wouldn't fit in aggregate GPU memory
  • Bigger batches: Use memory savings for increased batch sizes
  • Longer sequences: Handle extended context lengths
  • More devices: Scale to larger numbers of GPUs effectively

Communication Patterns

All-Gather Operations

  • Frequency: Parameter gathering before each layer computation
  • Size: Only required parameter subset
  • Optimization: Overlap with computation when possible

All-Reduce for Gradients

  • Stage 1 & 2: Traditional gradient synchronization
  • Stage 3: Reduced communication volume due to partitioning
  • Bucketing: Efficient handling of small gradient groups

Trade-offs and Considerations

Communication Overhead

  • Increased frequency: More communication operations per training step
  • Network sensitivity: Performance heavily dependent on interconnect bandwidth
  • Latency impact: Higher communication latency affects training speed

Implementation Complexity

  • Stage progression: Each stage adds implementation complexity
  • Memory management: Sophisticated dynamic allocation required
  • Debugging difficulty: Distributed state makes debugging challenging

Framework Integration

DeepSpeed Implementation

  • Native support: ZeRO is core feature of Microsoft's DeepSpeed
  • Automatic optimization: Framework handles communication scheduling
  • Configuration: Simple parameter selection for different stages

Other Framework Support

  • PyTorch FSDP: Similar concepts in Fully Sharded Data Parallel
  • FairScale: Facebook's implementation of sharding strategies
  • Custom implementations: Framework-agnostic manual implementation possible

Performance Optimization

Stage Selection Strategy

Choose optimal stage based on:

  • Memory pressure: How severely memory constrained
  • Network bandwidth: Available inter-device communication
  • Model size: Larger models benefit more from aggressive stages
  • Batch size requirements: Memory needs for target batch size

Hybrid Approaches

  • Selective partitioning: Partition only specific components
  • Gradient accumulation: Combine with micro-batching strategies
  • Mixed precision: Coordinate with FP16/BF16 optimizations

Advanced Optimizations

ZeRO-Offload

  • CPU offloading: Move optimizer states to CPU memory
  • Heterogeneous memory: Utilize both GPU and CPU memory hierarchies
  • Bandwidth management: Balance GPU-CPU transfer costs

ZeRO-Infinity

  • NVMe integration: Use high-speed storage for parameter swapping
  • Memory hierarchy: GPU → CPU → NVMe memory management
  • Extremely large models: Train models larger than total system memory

See also