Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Activation Recomputation
page dédiée →Memory optimization technique that trades computation for memory by recomputing forward pass activations during the backward pass instead of storing them throughout training. Also known as gradient checkpointing, this technique is fundamental to scaling neural network training to larger models and batch sizes.
Core Concept
Fundamental Trade-off
- Memory Savings: 50-90% reduction in activation memory usage
- Computational Cost: 15-20% increase in total computation
- Net Benefit: Enables training larger models or batch sizes that wouldn't fit in memory otherwise
Why It Works
During standard training:
- Forward pass stores all intermediate activations
- Backward pass uses stored activations to compute gradients
- Peak memory occurs when all activations are stored simultaneously
With activation recomputation:
- Forward pass stores only selected checkpoint activations
- Backward pass recomputes needed activations from checkpoints
- Peak memory reduced to checkpoint storage plus recomputation working memory
Implementation Strategy
Checkpointing Approach
- Checkpoint Selection: Store activations at strategic layer boundaries
- Segment Recomputation: Recompute activations within segments during backprop
- Granularity Control: Balance checkpoint frequency vs. recomputation overhead
Typical Checkpoint Placement
For transformer models:
- Checkpoint at attention block boundaries
- Store attention outputs and feed-forward outputs
- Recompute internal attention and FFN activations as needed
Memory Calculation Example
For Llama 3 8B model:
- Standard Training: ~61.09 GB activation memory
- With Recomputation: ~6-30 GB activation memory (depending on checkpoint frequency)
- Total Savings: 50-90% activation memory reduction
Integration with Training Step Anatomy
Forward Pass Modifications
- Store only checkpoint activations instead of all intermediate results
- Continue normal forward computation but discard non-checkpoint activations
- Mark checkpoint boundaries for backward pass reference
Backward Pass Modifications
- When gradient computation needs missing activation:
- Locate nearest stored checkpoint
- Recompute forward pass from checkpoint to needed activation
- Use recomputed activation for gradient calculation
- Discard recomputed activation after use
Memory Dynamic Changes
Changes the typical training-step-anatomy memory patterns:
- Forward Pass: Lower peak due to limited activation storage
- Backward Pass: Micro-spikes during recomputation phases
- Overall: Significantly reduced memory footprint
Advanced Optimization Techniques
Selective Recomputation
Not all activations need recomputation:
- Cheap Operations: Always recompute (element-wise operations, layer norms)
- Expensive Operations: Consider checkpointing (attention, large matrix multiplications)
- Memory-Heavy: Prioritize for recomputation (large activation tensors)
Overlapping Strategies
- Computation-Communication Overlap: Recompute activations while communicating gradients
- Pipeline Integration: Coordinate recomputation with pipeline parallel stages
- Memory Pool Management: Efficiently manage temporary memory for recomputation
Hardware-Specific Tuning
- GPU Memory Hierarchy: Utilize L2 cache for frequently recomputed activations
- Tensor Core Optimization: Ensure recomputed operations use optimal data layouts
- Mixed Precision: Apply appropriate precision for recomputed vs. stored activations
Production Considerations
Implementation Frameworks
- PyTorch: Built-in
torch.utils.checkpointfunctionality - Nanotron: Production implementation used at Hugging Face
- Picotron: Educational reference implementations
Configuration Parameters
- Checkpoint Frequency: How often to store activations
- Recomputation Granularity: Size of recomputed segments
- Memory Budget: Target memory usage vs. compute overhead
Monitoring and Debugging
- Track recomputation overhead in training metrics
- Monitor memory usage patterns during recomputation phases
- Profile backward pass timing to optimize checkpoint placement
Distributed Training Integration
Multi-GPU Coordination
- Coordinate checkpoint placement across tensor parallel ranks
- Ensure recomputation doesn't create communication bottlenecks
- Balance memory savings vs. increased computation across devices
Pipeline Parallelism Interaction
- Coordinate recomputation with pipeline stage boundaries
- Optimize bubble time during recomputation phases
- Balance checkpoint storage across pipeline stages
Expert Parallelism Considerations
- Apply recomputation selectively to expert vs. shared layers
- Coordinate expert routing with recomputation scheduling
- Optimize memory usage across expert parallel groups
Mathematical Foundation
Memory Reduction Formula
For L layers with checkpoint every C layers:
- Standard Memory: O(L × batch_size × sequence_length × hidden_dim)
- With Checkpointing: O((L/C + C) × batch_size × sequence_length × hidden_dim)
- Optimal C: √L for balanced memory-computation trade-off
Computational
bfloat16 Precision
page dédiée →Brain floating-point 16-bit format designed to accelerate machine learning training while maintaining numerical stability. Provides significant speed and memory improvements over full float32 precision with minimal impact on model quality.
Technical Design
Format Structure:
- 16 bits total: 1 sign bit, 8 exponent bits, 7 mantissa bits
- Same exponent range as float32 (better than float16)
- Reduced mantissa precision compared to float32
- Direct truncation from float32 (no complex conversion)
Key Advantages:
- 2x memory reduction compared to float32
- Faster computation on modern GPUs (nvidia-b300, etc.)
- Better numerical stability than float16
- Seamless integration with existing training pipelines
Practical Applications
LLM Training Optimization:
- Used in gpu-mode-paris-2026 hackathon for 10-minute training constraints
- Enables larger models or batch sizes within memory limits
- Combined with gradient-accumulation for memory-efficient training
- Critical for competitive training scenarios requiring maximum speed
Memory Efficiency:
- Halves memory footprint for model weights and activations
- Enables training larger models on fixed hardware
- Reduces data transfer overhead in distributed-training
- Particularly effective with modern GPU architectures
Implementation Considerations
Training Pipeline Integration:
- Automatic mixed precision (AMP) frameworks handle conversions
- Master weights maintained in float32 for stability
- Gradients computed and accumulated in reduced precision
- Loss scaling prevents gradient underflow
Numerical Stability:
- Generally stable for most deep learning applications
- May require careful tuning for sensitive operations
- Loss scaling essential for gradient preservation
- Model-specific validation recommended
Performance Impact
Speed Improvements:
- Significant acceleration on tensor processing units
- Reduced memory bandwidth requirements
- Faster inter-GPU communication in distributed setups
- Essential for time-constrained training scenarios
Quality Trade-offs:
- Minimal impact on final model performance for most applications
- Occasional need for selective float32 operations
- Benefits typically outweigh precision costs
- Critical enabler for large-scale training
See also
Distributed Training
page dédiée →Training neural networks across multiple GPUs or machines to enable larger models, bigger batch sizes, and faster training. Essential for modern LLM development and production ML systems, representing the foundation for ultra-scale AI development.
Core Training Process
Model Replication: Identical copy of model parameters instantiated on each GPU, ensuring synchronized starting points.
Data Parallelism: Each GPU processes different subset of the training batch, maximizing parallel computation efficiency.
Gradient Synchronization: After backward pass, all GPUs exchange and average their computed gradients before parameter updates.
Parameter Updates: Each GPU applies the averaged gradients to its local model copy, maintaining synchronization across all devices.
Technical Implementation
Distributed Data Parallel (DDP)
PyTorch Implementation: Standard approach using torch.nn.parallel.DistributedDataParallel for automatic gradient synchronization.
Communication Backend: Uses NCCL (NVIDIA Collective Communication Library) for high-performance GPU-to-GPU communication.
Process Groups: Each GPU runs in separate process, coordinating through inter-process communication protocols.
Practical Example (32-GPU Setup)
# Each GPU processes portion of global batch
global_batch_size = 128 * 32 # 4096 total examples
local_batch_size = 128 # 128 examples per GPU
# Training loop
for batch in dataloader:
loss = model(batch) # Local computation
loss.backward() # Local gradient computation
# Automatic gradient averaging across all GPUs
optimizer.step() # Synchronized parameter update
Performance Characteristics
Linear Scaling: Ideal case achieves N× speedup with N GPUs, though communication overhead creates practical limitations.
Memory Efficiency: Combines with gradient-accumulation to simulate larger effective batch sizes without proportional memory increase.
Precision Optimization: Often uses bfloat16 precision to reduce memory usage and increase training speed without significant accuracy loss.
Competition Context
In gpu-mode-paris-2026 hackathon:
- Hardware: 32× NVIDIA B300 GPUs in coordinated cluster
- Time Constraint: 10-minute training window maximizes importance of efficient parallelization
- Optimization Target: Achieve lowest validation loss through optimal distributed resource utilization
Communication Patterns
AllReduce Operations: Primary communication pattern for gradient averaging across all GPUs simultaneously.
Bandwidth Requirements: High-speed interconnects (InfiniBand, NVLink) critical for minimizing communication overhead.
Synchronization Overhead: Trade-off between communication frequency and gradient staleness affects overall training efficiency.
This represents the backbone technology enabling training of modern large language models, from GPT-3 to contemporary 100B+ parameter models that cannot fit on single machines.
See also
Edge AI Optimization
page dédiée →Specialized techniques for deploying AI models on resource-constrained edge devices, focusing on memory efficiency, latency optimization, and task-specific performance rather than general capabilities.
Core Constraints
Memory-Bound Operations
- Models must operate within strict memory limits (<3B parameters)
- Memory bandwidth more limiting than computational power
- Parameter efficiency critical for deployment success
Latency Requirements
- Sub-100ms response times required for user-facing applications
- Fast prefill more important than decode speed optimization
- Real-time inference constraints shape architecture decisions
Device-Specific Optimization
- Mobile processors (Galaxy S24 Ultra, Ryzen HX 370)
- CPU-optimized inference paths
- Hardware-specific quantization strategies (4-bit with llama.cpp)
Architecture Strategies
Parameter Distribution
- Optimize embedding layer size (19% vs traditional 63% allocation)
- Balance between knowledge storage and computational efficiency
- Effective model size through strategic parameter allocation
Operator Efficiency
- Gated Short Convolution blocks show 2.5x better cost ratios
- Replace attention mechanisms with more efficient alternatives
- Hardware-specific operator optimization (CPU vs GPU paths)
Model Size Targets
- <1GB models for on-device reasoning (LFM2.5-1.2B-Thinking)
- Sub-3B parameter counts for memory-bound constraints
- Task-specific models over general-purpose alternatives
Inference Optimization
CPU Optimization
- llama.cpp integration with 4-bit quantization
- Memory-efficient attention alternatives (ShortConv)
- Optimized operator cost ratios for CPU decode
GPU Batch Processing
- SGLang integration for concurrent inference
- Scaling performance with multiple simultaneous requests
- Input/output token optimization (1024/256 token targets)
Mobile Deployment
- On-device profiling and optimization
- Platform-specific performance tuning
- Battery and thermal management considerations
Training Considerations
Task-Specific Focus
- Narrow domain optimization over general capabilities
- Easy adaptation to new domain-specific data
- Post-training efficiency for specialized tasks
Memory-Aware Training
- Architecture choices informed by deployment constraints
- Parameter allocation strategies during training
- Inference-first design philosophy
Performance Metrics
Latency Benchmarks
- Sub-100ms response time requirements
- Prefill speed optimization priorities
- Real-time inference capability
Memory Efficiency
- Model size under deployment constraints
- Runtime memory usage optimization
- Quantization impact on accuracy vs efficiency
Throughput Scaling
- Concurrent request handling
- Batch processing optimization
- Resource utilization efficiency
See also
Expert Model Loading
page dédiée →Memory optimization technique demonstrated in apple-siri-architecture where specialized model components are dynamically loaded from storage into RAM on a per-query basis, enabling large-scale AI capabilities on memory-constrained devices.
Technical Approach
Dynamic Loading System
- NAND-to-RAM transfer: Experts stored in flash storage, loaded as needed
- Query-specific activation: Different specialists loaded based on request type
- Memory footprint optimization: Temporary loading reduces permanent RAM usage
- Response time trade-off: Storage access latency vs. memory conservation
Architecture Benefits
- Large model capacity: 20B-parameter model on mobile hardware
- Memory efficiency: Avoid permanent allocation of all model components
- Specialization: Different experts for different query types
- Scalability: Can support more experts than would fit in RAM
Implementation Challenges
Performance Considerations
- Loading latency: Time required to transfer experts from storage
- Storage wear: Frequent NAND access may impact device longevity
- Prediction accuracy: Must correctly anticipate which experts to load
- Caching strategy: Optimizing which experts remain in memory
Technical Requirements
- Fast storage: High-speed NAND access for reasonable response times
- Prediction models: Systems to determine expert requirements from queries
- Memory management: Efficient allocation and deallocation of expert models
- Error handling: Graceful fallbacks when expert loading fails
Mobile AI Innovation
Resource Constraint Solutions
Expert loading represents creative adaptation to mobile hardware limitations, enabling sophisticated AI without requiring massive RAM allocation.
Privacy Implications
On-device expert loading supports privacy-preserving AI by avoiding cloud-based processing for sensitive queries.
Broader Applications
Edge Computing
The technique could apply to other resource-constrained environments requiring sophisticated AI capabilities.
Cost Optimization
Cloud deployments might use similar approaches to optimize memory costs in serving infrastructure.
Future Developments
- Faster storage technologies: Reducing loading latency
- Better prediction models: More accurate expert selection
- Hybrid approaches: Combining on-device and cloud experts
- Cross-platform adaptation: Applying to other mobile and edge platforms
See also
- apple-siri-architecture
- On-Device AI
- Mobile Optimization
- Memory Management
Gemma 4 QAT
page dédiée →Quantization-Aware Training (QAT) implementation for Gemma 4 models that achieves significant memory reduction while preserving performance. Represents advancement in efficient model deployment for resource-constrained environments.
Performance Characteristics
Memory Reduction: ~4x less memory usage compared to standard Gemma 4 models while maintaining comparable performance.
Mobile Optimization: Gemma 4 E2B variant fits in approximately 1GB using specialized mobile quantization format.
Performance Preservation: QAT training maintains model capabilities despite aggressive quantization.
Technical Implementation
Quantization-Aware Training: Models trained with quantization effects incorporated during training phase, enabling better preservation of capabilities compared to post-training quantization.
Mobile Format: Specialized quantization format optimized for mobile and edge deployment scenarios.
Hardware Integration: Optimized for deployment across various hardware configurations with limited memory.
Integration Support
llama.cpp Compatibility: Gemma 4 MTP merged into llama.cpp for faster decoding when paired with QAT checkpoints.
Ecosystem Support: Broad compatibility with existing inference frameworks and serving infrastructure.
Impact
Enables deployment of advanced language models in previously infeasible environments, advancing democratization of AI capabilities through efficient resource utilization.
See also
- quantization
- mobile-deployment
- inference-optimization
Gradient Accumulation
page dédiée →Memory optimization technique that enables training with larger effective batch sizes without proportionally increasing memory usage. Achieves this by processing data in smaller micro-batches and accumulating their gradients before performing parameter updates.
Core Mechanism
Problem Solved: GPU memory limitations prevent training with optimal batch sizes (e.g., 128 examples requiring 65GB when GPU has only 40GB available).
Solution Approach: Split large batch into smaller micro-batches, process sequentially while accumulating gradients, then perform single parameter update with accumulated gradients.
# Traditional approach (memory overflow)
loss = model(batch_128_examples) # OOM Error
# Gradient accumulation approach
total_loss = 0
for micro_batch in split_in_4(batch_128_examples): # 32 examples each
loss = model(micro_batch) # 8.75GB per micro-batch
loss.backward() # Accumulate gradients
total_loss += loss
optimizer.step() # Update with accumulated gradients
Benefits of Larger Effective Batches
Gradient Stability: Averaging over more examples reduces gradient noise, leading to smoother optimization landscapes.
Better Generalization: Models see more data diversity per training step, improving ability to generalize to unseen data.
Faster Convergence: Fewer total optimization steps required to reach target performance due to more informative gradient estimates.
Implementation in Practice
Hyperparameter Configuration: Common configurations use 4-8 gradient accumulation steps, effectively multiplying batch size by that factor.
Memory vs. Compute Trade-off: Exchanges additional forward pass computations for reduced memory requirements, enabling training of larger models on constrained hardware.
Distributed Training Integration: Works synergistically with distributed data parallel (DDP) training, where each GPU accumulates gradients independently before cross-GPU synchronization.
Competition Applications
In the gpu-mode-paris-2026 hackathon context:
- Standard Configuration: 4 micro-steps accumulating to effective batch size of 128
- Memory Constraints: Enables training on NVIDIA B300 GPUs without memory overflow
- Performance Optimization: Critical for achieving competitive validation loss within 10-minute training window
This technique represents a foundational optimization in modern deep learning, used universally in training large language models like GPT, LLaMA, and other transformer architectures.
See also
Inference Optimization
page dédiée →The field of techniques and strategies to reduce computational cost, memory usage, and latency when running large transformer models in production. Critical for deploying powerful models at scale in real-world applications where cost and performance constraints must be balanced against model capability.
Fundamental Challenges
According to lilian-weng's analysis building on pope-et-al-2022, inference challenges stem from two primary factors beyond just increasing model size:
- memory-bandwidth-bottleneck: The rate at which data can be transferred between memory and processing units becomes the limiting factor
- autoregressive-generation: Sequential token generation prevents effective parallelization strategies
Core Optimization Strategies
Model Compression
- quantization: Reducing numerical precision of parameters and activations
- pruning: Removing less important parameters or connections
- knowledge-distillation: Training smaller student models to replicate larger teacher behavior
Architecture Optimization
- attention-optimization: Improving computational and memory efficiency of attention mechanisms
- Sparse attention patterns: Reducing quadratic scaling of attention computation
- Key-value caching: Optimizing memory access patterns in autoregressive generation
Hardware Optimization
- Mixed precision training: Leveraging different numerical precisions for different operations
- Memory layout optimization: Improving data access patterns
- Parallel processing strategies: Maximizing utilization of available compute resources
Production Considerations
Real-world deployment requires balancing multiple constraints:
- Latency requirements: Response time expectations
- Memory limitations: Available RAM and VRAM constraints
- Cost optimization: Computational expense vs. model capability
- Accuracy preservation: Maintaining model performance through optimization
Research Evolution
The field has evolved from simple model size reduction to sophisticated techniques that maintain model capability while dramatically reducing resource requirements. Current research focuses on finding optimal trade-offs between efficiency and performance.
See also
- memory-bandwidth-bottleneck
- autoregressive-generation
- model-compression
- lilian-weng
- pope-et-al-2022
LLM Scaling Techniques
page dédiée →Methods and strategies for training increasingly large language models, encompassing both model architecture scaling and training infrastructure scaling. Critical for developing state-of-the-art AI systems that require coordination across hundreds to thousands of GPUs.
Batch Size Scaling Evolution
The LLM training community has seen dramatic increases in batch sizes over time, reflecting improved distributed training capabilities:
- Llama 1: ~4M tokens per batch, 1.4 trillion total training tokens
- DeepSeek: ~60M tokens per batch, 14 trillion total training tokens
- DeepSeek-V3/R1: Dynamic scaling from 3,072 to 15,360 input sequences during initial 469B tokens, then maintained at 15,360
Token-Based Measurement
Modern LLM training reports batch sizes in tokens rather than samples to maintain independence from sequence length:
batch_size_tokens = batch_size_samples × sequence_length
This standardization enables consistent comparison across different model architectures and training configurations.
Three-Challenge Framework
All scaling techniques address one or more of three fundamental challenges identified in the ultra-scale-playbook:
1. Memory Usage (Hard Constraint)
- Nature: If training step doesn't fit in memory, training cannot proceed
- Components: Model weights, gradients, optimizer states, activations
- Solutions: memory-optimization, activation-recomputation, gradient accumulation
2. Compute Efficiency
- Goal: Maximize hardware utilization by reducing idle time
- Challenges: Data transfer overhead, GPU synchronization delays
- Solutions: Kernel fusion, mixed precision, optimized data pipelines
3. Communication Overhead
- Impact: Inter-GPU communication keeps hardware idle
- Strategy: Optimize intra-node (fast) vs inter-node (slower) bandwidth usage
- Approach: Overlap communication with computation whenever possible
Five-Dimensional Parallelism
Modern LLM training requires coordination across multiple parallelization dimensions:
- Data Parallelism: Distribute training data across GPUs
- Tensor Parallelism: Split model tensors across devices
- Pipeline Parallelism: Distribute model layers across pipeline stages
- Context Parallelism: Handle sequences longer than single-device memory
- Expert Parallelism: Distribute mixture-of-experts layers
Trade-off Optimization
Scaling techniques often involve trading one resource against another:
- Computation vs Memory: activation-recomputation trades 15-20% additional compute for 50-90% memory reduction
- Communication vs Parallelism: Higher parallelism increases communication overhead
- Batch Size vs Convergence: Larger batches improve hardware utilization but may affect convergence properties
Empirical Research Foundation
The ultra-scale-playbook provides data-driven optimization through over 4,000 scaling experiments:
- Up to 512 GPUs tested
- Systematic measurement of throughput and GPU utilization
- Normalized performance metrics across different model sizes
- Reproducible benchmarking methodology
Batch Size Sensitivity and Optimization
Convergence Considerations
Batch size affects model convergence in complex ways:
- Small batches: Quick initial progress but noisy gradients prevent optimal final performance
- Large batches: Accurate gradients but potentially slower convergence and compute waste
- Sweet spot: Typically 4-60 million tokens per batch for recent LLMs
Dynamic Scaling Strategies
Advanced training runs employ dynamic batch size scaling:
- Start with smaller batches for rapid initial convergence
- Gradually increase batch size as training progresses
- Balance convergence speed with computational efficiency
Memory-First Design Philosophy
Modern scaling techniques prioritize memory optimization since memory represents a hard constraint unlike compute or communication which can be optimized without preventing training:
- Memory prediction tools for capacity planning
- Systematic memory component analysis
- training-step-anatomy understanding for optimization
- Empirical memory profiling for validation
See also
- ultra-scale-playbook
- gpu-cluster-training
- memory-optimization
- Five-Dimensional Parallelism
- training-step-anatomy
- distributed-training
Memory Optimization
page dédiée →Techniques and strategies for managing GPU memory efficiently during neural network training, particularly for large language models where memory constraints often represent the primary bottleneck in scaling training to larger models and batch sizes.
Memory as Hard Constraint
Memory represents a hard constraint in LLM training - if a single training step doesn't fit in GPU memory, training simply cannot proceed. This makes memory optimization the most critical aspect of distributed training, as established in the ultra-scale-playbook.
Unlike compute or communication which can be optimized for efficiency, memory has absolute limits that must be respected for training to function at all.
Four Memory Components
GPU memory during training consists of four primary components:
1. Model Weights
- Neural network parameters
- Stored in various precisions (FP32, BF16, FP8)
- Size determined by model architecture
2. Gradients
- Computed during backward pass
- Typically same size as model weights
- Temporarily stored before optimization step
3. Optimizer States
- Often the largest component
- Adam optimizer stores momentum and variance (2x model size)
- Can dominate total memory usage
4. Activations
- Intermediate values from forward pass
- Needed for gradient computation during backward pass
- Can be traded for computation through activation-recomputation
Memory Profiling and Analysis
Empirical Measurement Approach
The ultra-scale-playbook emphasizes empirical memory profiling over theoretical calculations:
- PyTorch Memory Profiler: Step-by-step memory allocation tracking
- Dynamic Patterns: Memory usage varies significantly during training steps
- First Step Anomaly: Initial step shows different patterns due to caching allocator preparation
Training Step Memory Anatomy
- Forward Pass: Activations build up progressively
- Backward Pass: Gradients accumulate while activations are cleared
- Optimization: All gradients needed, optimizer states updated
Memory Optimization Techniques
Precision Management
- Mixed Precision Training: Use BF16/FP16 instead of FP32
- FP8 Training: Cutting-edge precision for maximum memory savings
- Gradient Scaling: Maintain numerical stability with lower precision
Activation Management
- activation-recomputation: Trade computation for memory (50-90% memory reduction)
- Gradient Checkpointing: Strategic activation storage points
- Progressive Clearing: Clear activations as soon as gradients computed
Optimizer Optimization
- zero-optimizer: Partition optimizer states across devices
- AdamW vs Adam: More memory-efficient optimizer variants
- State Precision: Lower precision for optimizer states
Batch Size Management
- gradient-accumulation: Simulate larger batches without memory increase
- Micro-batching: Process smaller chunks within larger logical batches
- Dynamic Batching: Adjust batch size based on sequence length
Memory Prediction and Tools
Theoretical Calculation
Memory usage can be estimated from:
- Tensor shapes (batch size, sequence length, hidden dimensions)
- Precision formats (4 bytes for FP32, 2 for BF16, 1 for FP8)
- Model architecture parameters
Empirical Tools
- Memory Prediction Tools: Hugging Face's memory estimation widgets
- Profiling Dashboards: Real-time memory usage visualization
- Benchmarking Suites: Systematic memory usage measurement
Memory Fragmentation and Allocation
PyTorch Caching Allocator
- Pre-allocates memory blocks to speed up subsequent allocations
- Causes first step anomaly in memory patterns
- Can lead to fragmentation reducing usable memory
CUDA Kernel Overhead
- Kernels typically require 1-2 GB of GPU memory
- Constant overhead independent of model size
- Must be factored into memory budget
Trade-offs and Strategies
Computation-Memory Trade-offs
- Recomputation: Use more computation to reduce memory storage
- Batching: Larger batches improve efficiency but increase memory
- Precision: Lower precision saves memory but may affect convergence
Memory-Communication Balance
- Smaller models per device reduce memory but increase communication
- Optimal balance depends on interconnect bandwidth
- Different strategies for intra-node vs inter-node communication
Scaling Implications
Memory optimization becomes increasingly critical at scale:
- Single GPU: Focus on activation recomputation and mixed precision
- Multi-GPU: Add optimizer state sharding and gradient compression
- Ultra-Scale: Combine all techniques with sophisticated parallelism strategies
Understanding memory patterns through empirical profiling enables informed decisions about which optimization techniques to apply for specific training configurations.
See also
- training-step-anatomy
- activation-recomputation
- gradient-accumulation
- zero-optimizer
- Mixed Precision Training
- ultra-scale-playbook
- gpu-cluster-training
Model Compression
page dédiée →The field of techniques for reducing the size and computational requirements of neural networks while maintaining performance. Critical for deploying large models in resource-constrained environments and reducing inference costs in production systems.
Technical Foundation
As systematically analyzed by lilian-weng, model compression is essential for addressing inference-optimization challenges, particularly the memory-bandwidth-bottleneck and constraints imposed by autoregressive-generation in large transformer models.
Core Compression Techniques
quantization
Reducing numerical precision of model parameters and activations:
- 8-bit quantization: ~4x size reduction with minimal accuracy loss
- 4-bit quantization: ~8x size reduction requiring careful implementation
- Mixed precision: Balancing compression with accuracy preservation
- Directly addresses memory bandwidth constraints by reducing data transfer requirements
pruning
Removing less important model components:
- Unstructured pruning: Removing individual parameters based on magnitude or importance
- Structured pruning: Removing entire neurons, channels, or blocks
- Sparse models: Maintaining connectivity patterns while reducing active parameters
- Enables hardware acceleration through specialized sparse computation
knowledge-distillation
Training smaller models to replicate larger model behavior:
- Teacher-student framework: Large model guides smaller model training
- Soft targets: Using probability distributions rather than hard classifications
- Feature matching: Aligning intermediate representations between models
- Enables deployment of powerful model capabilities in constrained environments
Optimization Objectives
Memory Efficiency
- Reducing model size for storage and RAM requirements
- Enabling deployment on edge devices with limited memory
- Addressing memory bandwidth bottlenecks in inference
Computational Efficiency
- Reducing FLOPs required for inference
- Improving throughput and reducing latency
- Enabling real-time applications with strict timing constraints
Energy Efficiency
- Reducing power consumption for mobile deployment
- Extending battery life in portable devices
- Reducing operational costs in data center deployment
Implementation Strategies
Progressive Compression
- Gradually applying compression techniques to monitor accuracy impact
- Starting with less aggressive settings and increasing compression
- Allows finding optimal points in accuracy-efficiency trade-off space
Multi-technique Combination
- Applying quantization, pruning, and distillation together
- Techniques can be complementary when properly orchestrated
- Requires careful coordination to avoid compounding accuracy losses
Hardware-Aware Compression
- Tailoring compression to target deployment hardware
- Leveraging hardware-specific optimizations and constraints
- Ensuring compressed models can efficiently utilize available resources
Challenges and Trade-offs
Accuracy Preservation
- Maintaining model performance while reducing complexity
- Different tasks and architectures have varying compression tolerance
- Requires careful evaluation and validation processes
Hardware Compatibility
- Ensuring compressed models work efficiently on target hardware
- Software framework support for optimized compressed model formats
- Balancing theoretical compression with practical deployment benefits
Development Complexity
- Additional engineering effort to implement and validate compression
- Need for specialized tools and frameworks
- Increased testing and validation requirements
See also
QAT Quantization
page dédiée →Quantization-Aware Training (QAT) is an advanced optimization technique that incorporates quantization effects during the training process rather than applying quantization post-hoc. This approach enables dramatic memory reduction while preserving model performance, as demonstrated by gemma-4's achievement of ~4x memory reduction.
Core Methodology
Training Integration: Quantization effects are simulated during training, allowing the model to learn to compensate for precision loss inherent in quantized representations.
Precision Optimization: Balances numerical precision with memory efficiency by training weights to be robust to quantization errors.
Hardware Awareness: Training process considers target deployment hardware constraints, optimizing for specific memory and computational limitations.
Performance Characteristics
Memory Efficiency: Achieves approximately 75% memory reduction compared to full-precision models while maintaining competitive performance metrics.
Quality Preservation: Unlike post-training quantization, QAT maintains model capabilities by learning to work within quantization constraints during training.
Deployment Viability: Enables deployment in severely resource-constrained environments where full-precision models would be infeasible.
Implementation Benefits
Mobile Deployment: Makes advanced AI models viable for smartphone and edge device deployment with ~1GB memory footprints.
Infrastructure Costs: Dramatically reduces serving costs and hardware requirements for model deployment at scale.
Accessibility: Democratizes access to advanced AI capabilities in environments with strict resource constraints.
Technical Challenges
Training Complexity: Requires careful tuning of quantization parameters and training schedules to achieve optimal performance.
Hardware Specificity: Optimal quantization strategies may vary significantly across different target deployment hardware.
Validation Requirements: Extensive testing needed to ensure quantized models maintain acceptable performance across diverse use cases.
Industry Impact
Edge AI Revolution: Enables practical deployment of sophisticated models on consumer devices and IoT hardware.
Cost Optimization: Provides pathway for significant infrastructure cost reduction in large-scale AI deployments.
Innovation Catalyst: Opens new possibilities for AI applications in resource-constrained environments previously considered infeasible.
See also
- gemma-4
- Edge AI
- quantization
- mobile-ai
- Efficient Inference
Quantization
page dédiée →The process of reducing the numerical precision of neural network parameters and activations from higher precision formats (like FP32 or FP16) to lower precision formats (like INT8 or INT4). A critical technique in model-compression for reducing memory usage, improving inference speed, and enabling deployment on resource-constrained hardware.
Technical Foundation
As analyzed by lilian-weng, quantization is one of the core strategies in inference-optimization for addressing the memory-bandwidth-bottleneck that constrains large transformer model deployment.
Types of Quantization
Post-Training Quantization (PTQ)
- Applied to already-trained models without additional training
- Faster to implement but may have larger accuracy drops
- Suitable for models with sufficient redundancy
Quantization-Aware Training (QAT)
- Incorporates quantization simulation during training process
- Better accuracy preservation but requires more computational resources
- Model learns to be robust to quantization effects
Precision Levels
8-bit (INT8)
- Reduces model size by ~4x compared to FP32
- Generally maintains good accuracy with proper calibration
- Well-supported across hardware platforms
4-bit (INT4)
- Aggressive compression reducing size by ~8x
- Requires careful implementation to maintain accuracy
- Increasingly supported in modern inference frameworks
Mixed Precision
- Different layers or operations use different precisions
- Balances compression with accuracy preservation
- Allows fine-tuning of the precision-accuracy trade-off
Implementation Considerations
Calibration Dataset
- Representative data used to determine quantization parameters
- Critical for maintaining model accuracy
- Should reflect actual deployment data distribution
Quantization Schemes
- Symmetric: Zero point is at the center of the range
- Asymmetric: Zero point can be offset for better range utilization
- Per-channel vs. per-tensor: Granularity of quantization parameters
Memory and Performance Benefits
Memory Reduction
- Direct reduction in model size proportional to precision decrease
- Enables deployment on resource-constrained devices
- Addresses memory-bandwidth-bottleneck by reducing data transfer requirements
Speed Improvements
- Lower precision arithmetic can be computed faster
- Hardware-specific optimizations for quantized operations
- Reduced memory access time due to smaller data sizes
Energy Efficiency
- Lower precision operations consume less energy
- Particularly important for edge deployment
- Extends battery life in mobile applications
Challenges and Limitations
Accuracy Degradation
- Some accuracy loss is typically unavoidable
- Certain model architectures more sensitive to quantization
- Requires careful evaluation of accuracy-efficiency trade-offs
Hardware Support
- Not all hardware platforms support all quantization schemes
- Software frameworks may have varying levels of optimization
- Need to match quantization approach to deployment target
See also
SWA Attention
page dédiée →Sliding Window Attention mechanism implemented in gemma-3-270m architecture as part of modern efficiency-focused attention innovations. Provides memory-efficient attention computation by limiting the attention window to a fixed size rather than full sequence attention.
Architecture Implementation
Gemma 3 270M Configuration
- Layer structure: 24 layers total
- Ratio: 3:1 GDN/Gated Attention with SWA
- Parameter allocation: 63% of total model parameters
- Effective size: ~100M parameters despite 270M total
Sliding Window Mechanism
Memory Efficiency
- Limits attention computation to a fixed window size
- Reduces memory requirements for long sequences
- Maintains performance while decreasing computational complexity
- Enables processing of longer contexts with constrained resources
Computational Benefits
- Linear memory scaling instead of quadratic with sequence length
- Improved cache efficiency for edge deployment
- Reduced latency for real-time applications
Performance Analysis
CPU Inference Cost
In maxime-labonne's M4 Max CPU benchmarking, SWA demonstrates moderate computational overhead compared to shortconv's minimal cost profile but significantly better than traditional full attention mechanisms.
Edge Deployment Advantages
- Suitable for memory-constrained environments
- Predictable memory usage regardless of input length
- Fast prefill performance critical for edge applications
Comparison with Other Mechanisms
Efficiency Ranking (CPU decode cost)
- shortconv - Lowest cost, optimized for CPU
- SWA - Moderate cost, memory efficient
- GDN - Higher cost, multimodal capabilities
- GLA/GQA - Variable cost depending on configuration
Applications
Edge Model Deployment
Particularly effective for:
- Real-time text processing
- Memory-constrained devices
- Long document processing with limited resources
- Applications requiring predictable latency
See also
- gemma-3-270m
- shortconv
- attention-mechanism
- edge-models
Ultra-Scale Playbook
page dédiée →Comprehensive 240+ page methodology and knowledge base for scaling LLM training from single GPUs to thousands of coordinated GPUs, developed by hugging-face through systematic empirical research. Represents the first open-sourcing of previously proprietary distributed training knowledge held within elite industry labs.
Knowledge Democratization Mission
The Ultra-Scale Playbook addresses a critical gap in the AI training ecosystem: while foundation models are openly available, the knowledge and techniques for training them at scale remained "well kept within a handful of big industry labs." This comprehensive resource lifts the veil on distributed training methodologies that were previously proprietary.
Three-Pillar Foundation
The playbook is built on three complementary foundations:
1. Theoretical Understanding
- Quick introductions to concepts and methods
- High-level explanations of advantages and limitations
- Memory breakdown analysis for Transformer models
- Understanding of when and why memory constraints occur
2. Clear Code Implementations
- picotron: Educational implementations in single, self-contained files for learning
- nanotron: Production-ready codebase used at Hugging Face
- Theory-to-code translations revealing implementation details and edge cases
3. Real Training Efficiency Benchmarks
- Over 4,000 systematic scaling experiments
- Up to 512 GPU cluster configurations tested
- Infrastructure-specific optimization guidance
- Reproducible performance measurements
Three Core Challenges Framework
All distributed training techniques address one or more of these fundamental challenges:
- Memory Usage: Hard constraint - if a training step doesn't fit in memory, training cannot proceed
- Compute Efficiency: Maximizing hardware utilization, minimizing idle time and data transfer delays
- Communication Overhead: Reducing inter-GPU communication that keeps devices idle, optimizing bandwidth usage
Memory as Hard Constraint
The playbook establishes memory as the primary bottleneck in LLM training, consisting of four critical components:
- Model Weights: Parameters of the neural network
- Gradients: Computed during backward pass
- Optimizer States: Often the largest component (e.g., Adam momentum and variance)
- Activations: Intermediate values needed for gradient computation
Training Step Anatomy
Detailed analysis of what happens during a single training step:
- Forward Pass: Activations build up as inputs pass through layers
- Backward Pass: Gradients computed while activations progressively cleared
- Optimization Step: All gradients needed, optimizer states updated
First Step Anomaly
The first training step exhibits different memory patterns due to PyTorch caching allocator preparation work. This can lead to successful first steps followed by OOM failures in subsequent steps due to optimizer state buildup.
Batch Size Evolution
Modern LLM training uses token-based batch sizes for sequence-length independence:
- Llama 1: ~4M tokens per batch, 1.4 trillion total tokens
- DeepSeek: ~60M tokens per batch, 14 trillion total tokens
- Sweet Spot: 4-60 million tokens per batch for current LLM training
Educational Philosophy
Combines systematic empirical research with educational accessibility. The approach prioritizes understanding over optimization, making complex distributed training concepts accessible through clear explanations, focused implementations, and real-world benchmarking data.
Empirical Research Foundation
Built on systematic experimentation rather than purely theoretical analysis:
- Over 4,100 distributed experiments (16k+ including test runs)
- Systematic scanning of distributed training layouts and model sizes
- Infrastructure-specific optimization insights
- Reproducible benchmarking methodologies
Impact
Represents a fundamental shift in knowledge sharing for AI training, moving previously proprietary expertise into the open-source domain. Enables researchers and practitioners to understand and implement ultra-scale training without starting from scratch or reverse-engineering techniques from scattered papers.
See also
- memory-optimization
- training-step-anatomy
- distributed-training
- hugging-face
- gpu-cluster-training
- picotron
- nanotron
ZeRO Optimizer
page dédiée →Zero Redundancy Optimizer (ZeRO) is an advanced memory optimization technique for distributed training that eliminates memory redundancy by partitioning optimizer states, gradients, and parameters across devices while maintaining training efficiency.
Core Problem: Memory Redundancy
Traditional Data Parallelism Issues
In standard distributed training:
- Each GPU maintains complete copy of model parameters
- Each GPU stores full optimizer states (often 2-3x parameter size)
- Each GPU accumulates complete gradient set
- Result: Massive memory redundancy across devices
Memory Components
For a model with P parameters using Adam optimizer:
- Model parameters: P values
- Gradients: P values
- Optimizer states: 2P values (momentum + variance)
- Total per GPU: 4P values × number of GPUs
ZeRO Stages
Stage 1: Optimizer State Partitioning
- Partition: Optimizer states across devices
- Memory reduction: 4x reduction for Adam optimizer
- Communication: Gather required states during optimization
- Benefit: Significant memory savings with minimal overhead
Stage 2: Gradient Partitioning
- Partition: Gradients in addition to optimizer states
- Memory reduction: 8x reduction total
- Communication: All-reduce only assigned gradient partitions
- Synchronization: Gradients distributed and synchronized efficiently
Stage 3: Parameter Partitioning
- Partition: Model parameters across devices
- Memory reduction: Linear with number of devices
- Communication: Gather parameters as needed for forward/backward
- Complexity: Most aggressive but requires careful implementation
Implementation Strategy
Dynamic Parameter Management
Stage 3 requires sophisticated parameter handling:
- Forward pass: Gather required parameters just before computation
- Computation: Execute with temporarily assembled parameters
- Cleanup: Discard non-local parameters to free memory
- Backward pass: Repeat gathering for gradient computation
Communication Optimization
- Overlap: Hide parameter gathering with computation
- Prefetching: Anticipate parameter needs for next layers
- Bucketing: Group small parameters for efficient communication
Memory Efficiency Gains
Theoretical Reductions
For N devices:
- Stage 1: Memory per device = (P + P + 2P/N) = (2P + 2P/N)
- Stage 2: Memory per device = (P + P/N + 2P/N) = (P + 3P/N)
- Stage 3: Memory per device = (P/N + P/N + 2P/N) = 4P/N
Practical Benefits
- Larger models: Train models that wouldn't fit in aggregate GPU memory
- Bigger batches: Use memory savings for increased batch sizes
- Longer sequences: Handle extended context lengths
- More devices: Scale to larger numbers of GPUs effectively
Communication Patterns
All-Gather Operations
- Frequency: Parameter gathering before each layer computation
- Size: Only required parameter subset
- Optimization: Overlap with computation when possible
All-Reduce for Gradients
- Stage 1 & 2: Traditional gradient synchronization
- Stage 3: Reduced communication volume due to partitioning
- Bucketing: Efficient handling of small gradient groups
Trade-offs and Considerations
Communication Overhead
- Increased frequency: More communication operations per training step
- Network sensitivity: Performance heavily dependent on interconnect bandwidth
- Latency impact: Higher communication latency affects training speed
Implementation Complexity
- Stage progression: Each stage adds implementation complexity
- Memory management: Sophisticated dynamic allocation required
- Debugging difficulty: Distributed state makes debugging challenging
Framework Integration
DeepSpeed Implementation
- Native support: ZeRO is core feature of Microsoft's DeepSpeed
- Automatic optimization: Framework handles communication scheduling
- Configuration: Simple parameter selection for different stages
Other Framework Support
- PyTorch FSDP: Similar concepts in Fully Sharded Data Parallel
- FairScale: Facebook's implementation of sharding strategies
- Custom implementations: Framework-agnostic manual implementation possible
Performance Optimization
Stage Selection Strategy
Choose optimal stage based on:
- Memory pressure: How severely memory constrained
- Network bandwidth: Available inter-device communication
- Model size: Larger models benefit more from aggressive stages
- Batch size requirements: Memory needs for target batch size
Hybrid Approaches
- Selective partitioning: Partition only specific components
- Gradient accumulation: Combine with micro-batching strategies
- Mixed precision: Coordinate with FP16/BF16 optimizations
Advanced Optimizations
ZeRO-Offload
- CPU offloading: Move optimizer states to CPU memory
- Heterogeneous memory: Utilize both GPU and CPU memory hierarchies
- Bandwidth management: Balance GPU-CPU transfer costs
ZeRO-Infinity
- NVMe integration: Use high-speed storage for parameter swapping
- Memory hierarchy: GPU → CPU → NVMe memory management
- Extremely large models: Train models larger than total system memory
See also
- memory-optimization
- distributed-training
- [[DeepSpeed