~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

5D Parallelism

page dédiée →

Advanced distributed training methodology that simultaneously coordinates five dimensions of parallelism to enable ultra-scale LLM training across thousands of GPUs. Represents the state-of-the-art approach for training the largest language models by addressing different aspects of memory and computation scaling challenges.

Five Parallelism Dimensions

1. Data Parallelism: Distribute batch samples across GPUs

  • Replicates model on each GPU
  • Each GPU processes different data samples
  • Requires gradient synchronization after backward pass

2. Tensor Parallelism: Split model weights across GPUs

  • Distributes memory requirements for large models
  • Requires communication during forward/backward passes
  • Effective for memory-bound scenarios

3. Pipeline Parallelism: Distribute model layers across GPUs

  • Sequential processing with potential bubble overhead
  • Various scheduling schemes to minimize idle time
  • Enables training models larger than single-GPU memory

4. Context Parallelism: Distribute sequence processing (Ring Attention)

  • Handles sequences longer than single-GPU memory capacity
  • Specialized for attention computation optimization
  • Particularly valuable for long-context training

5. Expert Parallelism: Distribute experts in Mixture-of-Experts models

  • Leverages sparse activation patterns
  • Reduces per-GPU computation and memory requirements
  • Essential for efficient MoE training

Coordination Challenges

Multi-Dimensional Optimization: Each parallelism dimension addresses different scaling bottlenecks:

  • Memory constraints (tensor, pipeline parallelism)
  • Batch size scaling (data parallelism)
  • Sequence length limits (context parallelism)
  • Sparse computation efficiency (expert parallelism)

Communication Patterns: Different parallelism types require different communication patterns and timing, requiring careful coordination to maintain efficiency and correctness.

Load Balancing: Achieving optimal balance across all five dimensions simultaneously while maintaining high GPU utilization across the entire cluster.

Implementation Framework

The ultra-scale-playbook provides systematic methodology for configuring 5D parallelism:

Configuration Process:

  1. Memory Analysis: Determine which dimensions are needed to fit model in memory
  2. Batch Size Requirements: Configure data parallelism to achieve target global batch size
  3. Throughput Optimization: Balance all dimensions for maximum training efficiency
  4. Empirical Validation: Benchmark across configurations to find optimal settings

Scaling Benefits

Ultra-Scale Enablement: Makes training of models requiring thousands of GPUs practically feasible by distributing different aspects of the training workload across multiple parallelism axes.

Resource Utilization: Maximizes utilization of expensive GPU clusters by ensuring each parallelism dimension addresses its target bottleneck without redundancy.

Flexibility: Allows adaptation to different hardware configurations, model architectures, and training requirements through adjusting the balance across dimensions.

Hardware Considerations

Interconnect Optimization: Requires careful mapping of parallelism dimensions to hardware topology to optimize bandwidth usage for different communication patterns.

Memory Hierarchy: Different parallelism types stress different parts of the memory hierarchy (GPU memory, inter-GPU bandwidth, inter-node bandwidth).

Fault Tolerance: Complex coordination increases failure modes, requiring robust checkpointing and recovery strategies.

Modern Applications

State-of-the-Art Training: Used for training the largest current models that require coordination across hundreds to thousands of GPUs.

Cost Optimization: Enables efficient use of expensive GPU clusters by maximizing utilization through optimal parallelism configuration.

Research Democratization: The ultra-scale-playbook open-sources this methodology, making ultra-scale training accessible beyond elite industry labs.

See also

Activation Recomputation

page dédiée →

Memory optimization technique that trades computation for memory by recomputing forward pass activations during the backward pass instead of storing them throughout training. Also known as gradient checkpointing, this technique is fundamental to scaling neural network training to larger models and batch sizes.

Core Concept

Fundamental Trade-off

  • Memory Savings: 50-90% reduction in activation memory usage
  • Computational Cost: 15-20% increase in total computation
  • Net Benefit: Enables training larger models or batch sizes that wouldn't fit in memory otherwise

Why It Works

During standard training:

  1. Forward pass stores all intermediate activations
  2. Backward pass uses stored activations to compute gradients
  3. Peak memory occurs when all activations are stored simultaneously

With activation recomputation:

  1. Forward pass stores only selected checkpoint activations
  2. Backward pass recomputes needed activations from checkpoints
  3. Peak memory reduced to checkpoint storage plus recomputation working memory

Implementation Strategy

Checkpointing Approach

  • Checkpoint Selection: Store activations at strategic layer boundaries
  • Segment Recomputation: Recompute activations within segments during backprop
  • Granularity Control: Balance checkpoint frequency vs. recomputation overhead

Typical Checkpoint Placement

For transformer models:

  • Checkpoint at attention block boundaries
  • Store attention outputs and feed-forward outputs
  • Recompute internal attention and FFN activations as needed

Memory Calculation Example

For Llama 3 8B model:

  • Standard Training: ~61.09 GB activation memory
  • With Recomputation: ~6-30 GB activation memory (depending on checkpoint frequency)
  • Total Savings: 50-90% activation memory reduction

Integration with Training Step Anatomy

Forward Pass Modifications

  • Store only checkpoint activations instead of all intermediate results
  • Continue normal forward computation but discard non-checkpoint activations
  • Mark checkpoint boundaries for backward pass reference

Backward Pass Modifications

  • When gradient computation needs missing activation:
    1. Locate nearest stored checkpoint
    2. Recompute forward pass from checkpoint to needed activation
    3. Use recomputed activation for gradient calculation
    4. Discard recomputed activation after use

Memory Dynamic Changes

Changes the typical training-step-anatomy memory patterns:

  • Forward Pass: Lower peak due to limited activation storage
  • Backward Pass: Micro-spikes during recomputation phases
  • Overall: Significantly reduced memory footprint

Advanced Optimization Techniques

Selective Recomputation

Not all activations need recomputation:

  • Cheap Operations: Always recompute (element-wise operations, layer norms)
  • Expensive Operations: Consider checkpointing (attention, large matrix multiplications)
  • Memory-Heavy: Prioritize for recomputation (large activation tensors)

Overlapping Strategies

  • Computation-Communication Overlap: Recompute activations while communicating gradients
  • Pipeline Integration: Coordinate recomputation with pipeline parallel stages
  • Memory Pool Management: Efficiently manage temporary memory for recomputation

Hardware-Specific Tuning

  • GPU Memory Hierarchy: Utilize L2 cache for frequently recomputed activations
  • Tensor Core Optimization: Ensure recomputed operations use optimal data layouts
  • Mixed Precision: Apply appropriate precision for recomputed vs. stored activations

Production Considerations

Implementation Frameworks

  • PyTorch: Built-in torch.utils.checkpoint functionality
  • Nanotron: Production implementation used at Hugging Face
  • Picotron: Educational reference implementations

Configuration Parameters

  • Checkpoint Frequency: How often to store activations
  • Recomputation Granularity: Size of recomputed segments
  • Memory Budget: Target memory usage vs. compute overhead

Monitoring and Debugging

  • Track recomputation overhead in training metrics
  • Monitor memory usage patterns during recomputation phases
  • Profile backward pass timing to optimize checkpoint placement

Distributed Training Integration

Multi-GPU Coordination

  • Coordinate checkpoint placement across tensor parallel ranks
  • Ensure recomputation doesn't create communication bottlenecks
  • Balance memory savings vs. increased computation across devices

Pipeline Parallelism Interaction

  • Coordinate recomputation with pipeline stage boundaries
  • Optimize bubble time during recomputation phases
  • Balance checkpoint storage across pipeline stages

Expert Parallelism Considerations

  • Apply recomputation selectively to expert vs. shared layers
  • Coordinate expert routing with recomputation scheduling
  • Optimize memory usage across expert parallel groups

Mathematical Foundation

Memory Reduction Formula

For L layers with checkpoint every C layers:

  • Standard Memory: O(L × batch_size × sequence_length × hidden_dim)
  • With Checkpointing: O((L/C + C) × batch_size × sequence_length × hidden_dim)
  • Optimal C: √L for balanced memory-computation trade-off

Computational

bfloat16 Precision

page dédiée →

Brain floating-point 16-bit format designed to accelerate machine learning training while maintaining numerical stability. Provides significant speed and memory improvements over full float32 precision with minimal impact on model quality.

Technical Design

Format Structure:

  • 16 bits total: 1 sign bit, 8 exponent bits, 7 mantissa bits
  • Same exponent range as float32 (better than float16)
  • Reduced mantissa precision compared to float32
  • Direct truncation from float32 (no complex conversion)

Key Advantages:

  • 2x memory reduction compared to float32
  • Faster computation on modern GPUs (nvidia-b300, etc.)
  • Better numerical stability than float16
  • Seamless integration with existing training pipelines

Practical Applications

LLM Training Optimization:

  • Used in gpu-mode-paris-2026 hackathon for 10-minute training constraints
  • Enables larger models or batch sizes within memory limits
  • Combined with gradient-accumulation for memory-efficient training
  • Critical for competitive training scenarios requiring maximum speed

Memory Efficiency:

  • Halves memory footprint for model weights and activations
  • Enables training larger models on fixed hardware
  • Reduces data transfer overhead in distributed-training
  • Particularly effective with modern GPU architectures

Implementation Considerations

Training Pipeline Integration:

  • Automatic mixed precision (AMP) frameworks handle conversions
  • Master weights maintained in float32 for stability
  • Gradients computed and accumulated in reduced precision
  • Loss scaling prevents gradient underflow

Numerical Stability:

  • Generally stable for most deep learning applications
  • May require careful tuning for sensitive operations
  • Loss scaling essential for gradient preservation
  • Model-specific validation recommended

Performance Impact

Speed Improvements:

  • Significant acceleration on tensor processing units
  • Reduced memory bandwidth requirements
  • Faster inter-GPU communication in distributed setups
  • Essential for time-constrained training scenarios

Quality Trade-offs:

  • Minimal impact on final model performance for most applications
  • Occasional need for selective float32 operations
  • Benefits typically outweigh precision costs
  • Critical enabler for large-scale training

See also

Distributed Training

page dédiée →

Training neural networks across multiple GPUs or machines to enable larger models, bigger batch sizes, and faster training. Essential for modern LLM development and production ML systems, representing the foundation for ultra-scale AI development.

Core Training Process

Model Replication: Identical copy of model parameters instantiated on each GPU, ensuring synchronized starting points.

Data Parallelism: Each GPU processes different subset of the training batch, maximizing parallel computation efficiency.

Gradient Synchronization: After backward pass, all GPUs exchange and average their computed gradients before parameter updates.

Parameter Updates: Each GPU applies the averaged gradients to its local model copy, maintaining synchronization across all devices.

Technical Implementation

Distributed Data Parallel (DDP)

PyTorch Implementation: Standard approach using torch.nn.parallel.DistributedDataParallel for automatic gradient synchronization.

Communication Backend: Uses NCCL (NVIDIA Collective Communication Library) for high-performance GPU-to-GPU communication.

Process Groups: Each GPU runs in separate process, coordinating through inter-process communication protocols.

Practical Example (32-GPU Setup)

# Each GPU processes portion of global batch
global_batch_size = 128 * 32  # 4096 total examples
local_batch_size = 128        # 128 examples per GPU

# Training loop
for batch in dataloader:
    loss = model(batch)  # Local computation
    loss.backward()      # Local gradient computation
    # Automatic gradient averaging across all GPUs
    optimizer.step()     # Synchronized parameter update

Performance Characteristics

Linear Scaling: Ideal case achieves N× speedup with N GPUs, though communication overhead creates practical limitations.

Memory Efficiency: Combines with gradient-accumulation to simulate larger effective batch sizes without proportional memory increase.

Precision Optimization: Often uses bfloat16 precision to reduce memory usage and increase training speed without significant accuracy loss.

Competition Context

In gpu-mode-paris-2026 hackathon:

  • Hardware: 32× NVIDIA B300 GPUs in coordinated cluster
  • Time Constraint: 10-minute training window maximizes importance of efficient parallelization
  • Optimization Target: Achieve lowest validation loss through optimal distributed resource utilization

Communication Patterns

AllReduce Operations: Primary communication pattern for gradient averaging across all GPUs simultaneously.

Bandwidth Requirements: High-speed interconnects (InfiniBand, NVLink) critical for minimizing communication overhead.

Synchronization Overhead: Trade-off between communication frequency and gradient staleness affects overall training efficiency.

This represents the backbone technology enabling training of modern large language models, from GPT-3 to contemporary 100B+ parameter models that cannot fit on single machines.

See also

First Step Anomaly

page dédiée →

A well-documented phenomenon in PyTorch neural network training where the initial training step exhibits fundamentally different memory allocation patterns compared to all subsequent steps. This anomaly has critical implications for training reliability and memory planning, particularly in distributed training scenarios.

Anomalous Behavior Patterns

The first training step shows several distinct characteristics that differentiate it from subsequent steps:

Memory Allocation Plateau

  • Activation plateau: After initial rapid increase, activation memory plateaus for an extended period
  • Delayed clearing: Activations aren't cleared as quickly as in subsequent steps
  • Memory preparation: Extended periods of stable memory usage during initialization

Caching Allocator Preparation

The root cause of the anomaly lies in PyTorch's caching allocator behavior:

  • Memory block preparation: Pre-allocates memory blocks for future training steps
  • Search optimization: Eliminates need to search for free memory blocks in subsequent iterations
  • Performance front-loading: Optimization work done upfront to accelerate future steps

Critical Training Implications

OOM Failure Patterns

The first step anomaly creates a dangerous pattern where training appears successful initially but fails in subsequent steps:

  • Step 1 success: First step completes successfully due to delayed optimizer state buildup
  • Step 2 failure: Second step OOMs when optimizer states are fully established
  • False confidence: Initial success doesn't guarantee training sustainability

Memory Planning Challenges

  • Unreliable memory estimates: First step memory usage doesn't predict subsequent step requirements
  • Configuration validation: Cannot rely on first step completion for memory planning
  • Buffer requirements: Must plan for higher memory usage in subsequent steps

Optimizer State Buildup

A key component of the first step anomaly relates to optimizer state initialization:

Delayed State Creation

  • First step: Optimizer states begin building but aren't fully established
  • Subsequent steps: Full optimizer states (momentum, variance estimates) consume significant memory
  • Memory jump: Notable increase in memory usage between first and second steps

State Persistence

Once established, optimizer states persist throughout training:

  • Accumulated overhead: States accumulate across all model parameters
  • Memory multiplier: Can be 2-4× larger than model weights for optimizers like Adam
  • Persistent allocation: Memory doesn't fluctuate as much as activations

Diagnostic and Mitigation Strategies

Memory Profiling Considerations

When profiling training memory usage:

  • Multi-step profiling: Profile at least 3-4 steps to see true patterns
  • Exclude first step: Don't use first step memory as baseline for planning
  • Pattern recognition: Look for consistency starting from step 2

Training Configuration Validation

  • Conservative planning: Plan memory requirements based on steps 2+ rather than step 1
  • Gradual scaling: Test configurations with multiple steps before full training
  • Memory monitoring: Implement alerts for unexpected memory usage patterns

Production Considerations

  • Checkpoint timing: Don't checkpoint immediately after first step
  • Resource allocation: Size GPU memory based on steady-state requirements
  • Failure recovery: Design restart procedures that account for anomaly

Relationship to Other Memory Patterns

Interaction with Memory Components

The first step anomaly affects different memory components differently:

  • Activations: Most dramatically affected by the plateau behavior
  • Gradients: Build up normally but may persist longer
  • Optimizer states: Delayed initialization is core to the anomaly
  • Model weights: Unaffected by the anomaly

Distributed Training Impact

In distributed training scenarios, the first step anomaly can compound:

  • Synchronization delays: Different GPUs may show varying anomaly patterns
  • Communication overhead: Additional coordination required during initialization
  • Scaling challenges: Anomaly effects may amplify with larger GPU counts

Research and Development Implications

Algorithm Development

Understanding the first step anomaly is crucial for:

  • Memory optimization research: Accurate baseline measurements require accounting for anomaly
  • Training algorithm design: New techniques must consider initialization patterns
  • Benchmarking: Fair comparisons require excluding first step from performance metrics

Tool Development

  • Profiling tools: Should highlight first step anomaly in visualizations
  • Memory predictors: Must model anomaly separately from steady-state behavior
  • Training frameworks: Should warn users about anomalous first step behavior

See also

GPU Cluster Training

page dédiée →

Training large language models on clusters of hundreds to thousands of GPUs working in coordination. Represents the current frontier of LLM development where models require more compute than any single machine can provide.

Core Challenges

The ultra-scale-playbook identifies three fundamental challenges that all distributed training techniques must address:

1. Memory Usage (Hard Constraint)

  • Training cannot proceed if a single step exceeds GPU memory limits
  • Must coordinate memory usage across multiple memory types and locations
  • Optimizer states often consume more memory than model weights
  • See memory-optimization for specific techniques

2. Compute Efficiency

  • Hardware should spend maximum time computing vs waiting or transferring data
  • GPU utilization directly impacts training cost and completion time
  • Idle GPUs represent pure economic loss at cluster scale

3. Communication Overhead

  • Minimize time spent synchronizing between GPUs
  • Balance fast intra-node vs slower inter-node bandwidth
  • Overlap communication with computation whenever possible
  • Scale communication patterns efficiently as cluster size grows

Ultra-Scale Empirical Results

Hugging Face's research involved:

  • 4,000+ scaling experiments across different configurations
  • Up to 512 GPUs in coordinated training runs
  • 16,000+ total runs including testing and validation
  • Systematic measurement of throughput and GPU utilization

Results show that both throughput and utilization vary dramatically based on:

  • Model size and architecture
  • Parallelism strategy chosen
  • Hardware interconnect topology
  • Batch size and sequence length

Training Step Anatomy

Each training step in a cluster environment involves:

  1. Forward Pass: Data flows through distributed model components
  2. Backward Pass: Gradients computed and synchronized across devices
  3. Optimization: Parameter updates coordinated across the cluster

These steps become significantly more complex when model components are distributed across multiple devices using 5d-parallelism strategies.

Memory Components at Scale

The four memory components scale differently in cluster environments:

  • Model weights: Can be sharded across devices via tensor parallelism
  • Gradients: Must be synchronized, creating communication bottlenecks
  • Optimizer states: Largest component, benefits significantly from partitioning
  • Activations: Can be recomputed or distributed via sequence parallelism

Batch Size Scaling

Modern LLM training has seen dramatic batch size increases:

  • Llama 1: ~4M tokens per batch, 1.4T total training tokens
  • DeepSeek: ~60M tokens per batch, 14T total training tokens

This scaling requires sophisticated coordination of gradient-accumulation across hundreds of devices.

Implementation References

The playbook references two complementary codebases:

  • Picotron: Educational implementations for understanding concepts
  • Nanotron: Production-ready distributed training system used at Hugging Face

Hardware Considerations

Cluster training success depends heavily on:

  • GPU memory capacity: Determines maximum model sizes per device
  • Interconnect bandwidth: Affects communication-bound operations
  • Network topology: Influences optimal parallelism strategies
  • Memory bandwidth: Critical for activation and gradient handling

Profiling and Optimization

Successful cluster training requires systematic profiling to:

  • Identify memory usage patterns across devices
  • Measure communication overhead between nodes
  • Optimize kernel fusion opportunities
  • Balance different parallelism dimensions

See distributed-training-profiling for specific techniques.

See also

Gradient Accumulation

page dédiée →

Memory optimization technique that enables training with larger effective batch sizes without proportionally increasing memory usage. Achieves this by processing data in smaller micro-batches and accumulating their gradients before performing parameter updates.

Core Mechanism

Problem Solved: GPU memory limitations prevent training with optimal batch sizes (e.g., 128 examples requiring 65GB when GPU has only 40GB available).

Solution Approach: Split large batch into smaller micro-batches, process sequentially while accumulating gradients, then perform single parameter update with accumulated gradients.

# Traditional approach (memory overflow)
loss = model(batch_128_examples)  # OOM Error

# Gradient accumulation approach
total_loss = 0
for micro_batch in split_in_4(batch_128_examples):  # 32 examples each
    loss = model(micro_batch)  # 8.75GB per micro-batch
    loss.backward()  # Accumulate gradients
    total_loss += loss
optimizer.step()  # Update with accumulated gradients

Benefits of Larger Effective Batches

Gradient Stability: Averaging over more examples reduces gradient noise, leading to smoother optimization landscapes.

Better Generalization: Models see more data diversity per training step, improving ability to generalize to unseen data.

Faster Convergence: Fewer total optimization steps required to reach target performance due to more informative gradient estimates.

Implementation in Practice

Hyperparameter Configuration: Common configurations use 4-8 gradient accumulation steps, effectively multiplying batch size by that factor.

Memory vs. Compute Trade-off: Exchanges additional forward pass computations for reduced memory requirements, enabling training of larger models on constrained hardware.

Distributed Training Integration: Works synergistically with distributed data parallel (DDP) training, where each GPU accumulates gradients independently before cross-GPU synchronization.

Competition Applications

In the gpu-mode-paris-2026 hackathon context:

  • Standard Configuration: 4 micro-steps accumulating to effective batch size of 128
  • Memory Constraints: Enables training on NVIDIA B300 GPUs without memory overflow
  • Performance Optimization: Critical for achieving competitive validation loss within 10-minute training window

This technique represents a foundational optimization in modern deep learning, used universally in training large language models like GPT, LLaMA, and other transformer architectures.

See also

Knowledge Democratization

page dédiée →

The systematic effort to make previously proprietary or restricted technical knowledge publicly accessible, particularly in AI and machine learning where critical implementation details have traditionally been confined to elite industry laboratories.

Context in AI Training

Industry Knowledge Hoarding: While open-source models like Llama and DeepSeek are publicly available, the most challenging aspects of training these systems—coordination techniques for thousands of GPUs, distributed training optimizations, and scaling methodologies—have remained proprietary within major tech companies.

Access Asymmetry: Creates significant barriers between those with access to both models and training expertise versus those with only model access, limiting innovation and research capabilities across the broader community.

Hugging Face Ultra-Scale Playbook Example

Comprehensive Open-Sourcing: First systematic release of distributed training knowledge that was previously confined to elite industry labs, covering everything from theoretical concepts to production-ready implementations.

Empirical Research Sharing: Over 4,000 scaling experiments and benchmarking results made publicly available, providing data-driven insights that would typically require significant infrastructure investment to obtain.

Multi-Modal Knowledge Transfer: Combines theoretical explanations, practical code implementations, interactive tools, and real performance benchmarks to ensure knowledge is truly transferable.

Democratization Challenges

Implementation Complexity: Technical knowledge often requires significant expertise to apply effectively, meaning availability doesn't immediately translate to accessibility.

Infrastructure Requirements: Even with available knowledge, applying distributed training techniques still requires substantial computational resources.

Maintenance and Evolution: Keeping democratized knowledge current as techniques and hardware evolve requires ongoing community effort.

Impact on Research Ecosystem

Leveled Playing Field: Enables academic institutions, smaller companies, and independent researchers to apply state-of-the-art techniques previously available only to well-resourced organizations.

Innovation Acceleration: Broader access to advanced techniques can lead to faster innovation as more minds work on improving and extending the methodologies.

Educational Value: Creates opportunities for learning and skill development in areas that were previously inaccessible to most practitioners.

Community-Driven Extension

Discussion Platforms: Provision of community spaces for questions, feedback, and knowledge extension beyond the initial democratized resource.

Collaborative Improvement: Open-source approach allows community contributions to improve and extend the democratized knowledge base.

See also

Memory Optimization

page dédiée →

Techniques and strategies for managing GPU memory efficiently during neural network training, particularly for large language models where memory constraints often represent the primary bottleneck in scaling training to larger models and batch sizes.

Memory as Hard Constraint

Memory represents a hard constraint in LLM training - if a single training step doesn't fit in GPU memory, training simply cannot proceed. This makes memory optimization the most critical aspect of distributed training, as established in the ultra-scale-playbook.

Unlike compute or communication which can be optimized for efficiency, memory has absolute limits that must be respected for training to function at all.

Four Memory Components

GPU memory during training consists of four primary components:

1. Model Weights

  • Neural network parameters
  • Stored in various precisions (FP32, BF16, FP8)
  • Size determined by model architecture

2. Gradients

  • Computed during backward pass
  • Typically same size as model weights
  • Temporarily stored before optimization step

3. Optimizer States

  • Often the largest component
  • Adam optimizer stores momentum and variance (2x model size)
  • Can dominate total memory usage

4. Activations

  • Intermediate values from forward pass
  • Needed for gradient computation during backward pass
  • Can be traded for computation through activation-recomputation

Memory Profiling and Analysis

Empirical Measurement Approach

The ultra-scale-playbook emphasizes empirical memory profiling over theoretical calculations:

  • PyTorch Memory Profiler: Step-by-step memory allocation tracking
  • Dynamic Patterns: Memory usage varies significantly during training steps
  • First Step Anomaly: Initial step shows different patterns due to caching allocator preparation

Training Step Memory Anatomy

  1. Forward Pass: Activations build up progressively
  2. Backward Pass: Gradients accumulate while activations are cleared
  3. Optimization: All gradients needed, optimizer states updated

Memory Optimization Techniques

Precision Management

  • Mixed Precision Training: Use BF16/FP16 instead of FP32
  • FP8 Training: Cutting-edge precision for maximum memory savings
  • Gradient Scaling: Maintain numerical stability with lower precision

Activation Management

  • activation-recomputation: Trade computation for memory (50-90% memory reduction)
  • Gradient Checkpointing: Strategic activation storage points
  • Progressive Clearing: Clear activations as soon as gradients computed

Optimizer Optimization

  • zero-optimizer: Partition optimizer states across devices
  • AdamW vs Adam: More memory-efficient optimizer variants
  • State Precision: Lower precision for optimizer states

Batch Size Management

  • gradient-accumulation: Simulate larger batches without memory increase
  • Micro-batching: Process smaller chunks within larger logical batches
  • Dynamic Batching: Adjust batch size based on sequence length

Memory Prediction and Tools

Theoretical Calculation

Memory usage can be estimated from:

  • Tensor shapes (batch size, sequence length, hidden dimensions)
  • Precision formats (4 bytes for FP32, 2 for BF16, 1 for FP8)
  • Model architecture parameters

Empirical Tools

  • Memory Prediction Tools: Hugging Face's memory estimation widgets
  • Profiling Dashboards: Real-time memory usage visualization
  • Benchmarking Suites: Systematic memory usage measurement

Memory Fragmentation and Allocation

PyTorch Caching Allocator

  • Pre-allocates memory blocks to speed up subsequent allocations
  • Causes first step anomaly in memory patterns
  • Can lead to fragmentation reducing usable memory

CUDA Kernel Overhead

  • Kernels typically require 1-2 GB of GPU memory
  • Constant overhead independent of model size
  • Must be factored into memory budget

Trade-offs and Strategies

Computation-Memory Trade-offs

  • Recomputation: Use more computation to reduce memory storage
  • Batching: Larger batches improve efficiency but increase memory
  • Precision: Lower precision saves memory but may affect convergence

Memory-Communication Balance

  • Smaller models per device reduce memory but increase communication
  • Optimal balance depends on interconnect bandwidth
  • Different strategies for intra-node vs inter-node communication

Scaling Implications

Memory optimization becomes increasingly critical at scale:

  • Single GPU: Focus on activation recomputation and mixed precision
  • Multi-GPU: Add optimizer state sharding and gradient compression
  • Ultra-Scale: Combine all techniques with sophisticated parallelism strategies

Understanding memory patterns through empirical profiling enables informed decisions about which optimization techniques to apply for specific training configurations.

See also


Memory-Mapped Data Loading

page dédiée →

Data loading technique that maps files directly into virtual memory, allowing the operating system to handle data transfer efficiently without explicit file I/O operations. Essential for high-performance training scenarios where I/O bottlenecks can severely impact training speed.

Core Mechanism

Virtual Memory Mapping: Files are mapped directly into process address space, creating illusion that entire file contents are loaded in memory.

Operating System Optimization: OS handles actual data transfer from disk to RAM on-demand, using sophisticated caching and prefetching algorithms.

Reduced Memory Overhead: Only portions of file currently being accessed are loaded into physical RAM, enabling processing of datasets larger than available memory.

Advantages Over Traditional File I/O

Elimination of Double Buffering: Data flows directly from disk to model without intermediate copying steps.

OS-Level Caching: Operating system automatically caches frequently accessed data portions, improving subsequent access performance.

Concurrent Access: Multiple processes can efficiently share the same memory-mapped data without duplication.

Reduced System Calls: Eliminates repetitive read() operations that create kernel/user space transitions.

Implementation in Training Pipelines

# Traditional approach
with open('dataset.bin', 'rb') as f:
    data = f.read()  # Loads entire file into RAM

# Memory-mapped approach  
import mmap
with open('dataset.bin', 'rb') as f:
    mapped_data = mmap.mmap(f.fileno(), 0, access=mmap.ACCESS_READ)
    # Data accessed on-demand as training progresses

Competition Context

In gpu-mode-paris-2026 hackathon:

  • Pre-tokenized Data: Binary files containing processed tokens ready for model consumption
  • 10-Minute Constraint: I/O optimization critical when every second counts toward training efficiency
  • 32-GPU Scaling: Each process needs efficient access to shared dataset without memory duplication

Performance Characteristics

Startup Time: Near-instantaneous compared to traditional file loading approaches.

Memory Efficiency: Dataset size decoupled from physical RAM requirements.

Cache Locality: OS intelligently prefetches sequential data, optimizing for common training access patterns.

Scaling Benefits: Performance improvements compound with larger datasets and more complex training pipelines.

This technique represents fundamental infrastructure optimization in modern ML training, enabling efficient processing of multi-terabyte datasets that characterize contemporary language model development.

See also

Production-ready distributed training codebase developed and used by Hugging Face for training large language models at scale. Serves as the industrial-strength implementation of distributed training techniques described in the ultra-scale-playbook, designed for reliability and performance in production environments.

Production Design Philosophy

Nanotron embodies a production-first approach to distributed training, prioritizing reliability, performance, and maintainability over educational clarity:

Industrial Requirements

  • High reliability: Robust error handling and fault tolerance
  • Performance optimization: Optimized for throughput and efficiency
  • Scalability: Designed to handle hundreds to thousands of GPUs
  • Maintainability: Structured for long-term production use

Enterprise Features

  • Monitoring integration: Comprehensive metrics and logging
  • Checkpoint management: Robust model saving and recovery
  • Resource management: Efficient GPU and memory utilization
  • Configuration management: Flexible hyperparameter handling

Relationship to Educational Resources

Nanotron serves as the production counterpart to educational tools in Hugging Face's training ecosystem:

Complementary to Picotron

While picotron provides educational implementations, Nanotron offers:

  • Production optimization: Performance-tuned implementations
  • Enterprise features: Monitoring, logging, fault tolerance
  • Scalability focus: Designed for large-scale production training
  • Robustness: Battle-tested reliability for long training runs

Implementation of Ultra-Scale Playbook

Nanotron represents the practical application of ultra-scale-playbook principles:

  • Empirical validation: Real-world implementation of playbook techniques
  • Production testing: Validation through actual training workloads
  • Performance data: Source of benchmarking data used in playbook
  • Continuous improvement: Feedback loop between theory and practice

Technical Capabilities

Distributed Training Support

Nanotron implements the full spectrum of distributed training techniques:

Parallelism Strategies

PyTorch Memory Profiling

page dédiée →

Systematic methodology for analyzing GPU memory usage patterns during neural network training, essential for optimizing memory utilization and debugging out-of-memory issues in distributed training scenarios.

Memory Usage Patterns

Memory utilization exhibits dynamic behavior that varies significantly during training steps rather than remaining static:

Forward Pass: Rapid activation buildup as data flows through successive model layers, with memory usage increasing as intermediate results accumulate.

Backward Pass: Gradient accumulation phase where gradients build up while stored activations are progressively cleared as they're consumed for gradient computation.

Optimization Step: Peak memory usage period when all gradients and optimizer states are simultaneously present in memory before parameter updates.

Between Steps: Memory cleanup phase before starting next forward pass, though some persistent state remains from optimizer.

First Step Anomaly

The initial training step exhibits distinctly different memory patterns:

Caching Allocator Preparation: PyTorch's caching allocator performs significant setup work, preparing memory allocations to speed up subsequent steps by avoiding repeated memory block searches.

Activation Plateau: Unlike later steps, activations increase quickly then plateau for extended period during first step processing.

Optimizer State Initialization: Optimizer states appear after first step completion, offsetting baseline memory usage for all subsequent training.

OOM Implications: Common failure pattern where first step succeeds but subsequent steps cause out-of-memory errors due to optimizer state buildup and different allocation patterns.

Memory Components Analysis

Model Weights: Base parameter storage, typically smallest component in overall memory footprint.

Gradients: Accumulated during backward pass, usually similar in size to model weights but with different temporal patterns.

Optimizer States: Often largest memory consumer, especially with stateful optimizers like Adam that maintain momentum and variance estimates.

Activations: Most variable component, heavily dependent on batch size, sequence length, and model architecture depth.

Profiling Methodology

Empirical Measurement: Recommended over theoretical calculation due to complexity of predicting exact usage from model specifications alone.

Dynamic Analysis: Focus on understanding temporal patterns rather than static peak usage, as memory efficiency often depends on timing of allocations and deallocations.

Infrastructure Overhead: Account for additional memory requirements from CUDA kernels (typically 1-2GB), buffers, and fragmentation that affect available memory.

Practical Applications

Memory Planning: Enable accurate prediction of training requirements for different model and batch size configurations.

OOM Debugging: Identify specific phases of training causing memory pressure and potential optimization targets.

Configuration Optimization: Guide decisions on batch size, sequence length, and other hyperparameters based on memory constraints.

Distributed Training Design: Inform parallelization strategies by understanding memory bottlenecks and distribution opportunities.

Integration with Optimization Strategies

Memory profiling directly informs optimization techniques like activation-recomputation and gradient-accumulation by identifying which memory components offer the best reduction opportunities versus computational cost.

See also

PyTorch Profiling

page dédiée →

Performance and memory analysis tools within PyTorch for understanding resource utilization patterns during model training. Essential for optimizing distributed training configurations and diagnosing memory bottlenecks in large-scale LLM training.

Core Functionality

Memory Usage Tracking: Monitor dynamic memory allocation patterns throughout training steps, revealing how memory usage varies across forward pass, backward pass, and optimization phases.

Performance Analysis: Measure compute utilization, kernel execution times, and identify bottlenecks in training workflows.

Training Step Visualization: Generate detailed breakdowns of what happens during each phase of training, enabling optimization of memory-constrained configurations.

Memory Pattern Analysis

Dynamic Memory Tracking: Unlike static memory calculations, profiling reveals actual memory usage patterns:

  • Forward pass: Rapid activation buildup
  • Backward pass: Gradient accumulation with progressive activation cleanup
  • Optimization: Peak memory usage when gradients, parameters, and optimizer states coexist

First Step Anomaly Detection: Profiling reveals why initial training steps behave differently:

  • PyTorch caching allocator preparation during first step
  • Memory plateau during caching setup
  • Explains common OOM failures where first step succeeds but subsequent steps fail

Implementation Approaches

Basic Memory Tracking: Simple memory monitoring using torch.ones((1, 1)).to("cuda") to measure CUDA kernel overhead (typically 1-2GB baseline).

Comprehensive Profiling: Full PyTorch profiler integration for detailed analysis of training workflows and resource utilization patterns.

Production Monitoring: Continuous profiling during large-scale training to identify performance degradation or resource contention.

Distributed Training Applications

Multi-GPU Analysis: Profile memory and compute patterns across distributed training configurations to optimize resource allocation.

Communication Profiling: Identify communication bottlenecks and overlap opportunities between computation and data transfer.

Cluster Optimization: Use profiling data to optimize distributed training configurations across hundreds to thousands of GPUs.

Optimization Applications

Memory Budget Planning: Use profiling data to accurately predict memory requirements before scaling training to larger configurations.

Activation Recomputation Decisions: Profile memory vs. compute trade-offs to determine optimal recomputation strategies.

Batch Size Optimization: Analyze memory usage patterns to find optimal batch sizes within hardware constraints.

Debugging Common Issues

OOM Diagnosis: Identify exact causes of out-of-memory errors by analyzing memory usage patterns across training steps.

Performance Bottlenecks: Locate inefficient operations or memory allocation patterns that limit training throughput.

Resource Utilization: Understand whether training is memory-bound, compute-bound, or communication-bound.

See also

RSI Suppression

page dédiée →

Recursive Self-Improvement suppression mechanisms designed to limit AI models' effectiveness at accelerating their own development or creating more capable successor systems. anthropic's implementation in claude-fable 5 represents the first major deployment of silent-interventions specifically targeting AI research acceleration.

Implementation Details

claude-fable 5's RSI suppression operates through invisible modifications to model behavior, implemented via:

  • Prompt modification: Altering queries related to frontier AI development before processing
  • Steering vectors: Real-time adjustment of model representations during inference
  • Parameter-efficient fine-tuning (PEFT): Dynamic weight modifications targeting specific capabilities
  • Output degradation: Reducing quality of responses on targeted topics without user notification

Targeted Activities

The suppression mechanisms specifically target requests involving:

  • Building pretraining pipelines
  • Distributed training infrastructure design
  • ML accelerator architecture development
  • Model optimization and scaling techniques
  • Competing model development assistance

Scope and Statistics

According to anthropic's estimates:

  • Affects approximately 0.03% of total traffic
  • Concentrated in fewer than 0.1% of organizations
  • Does not affect "the vast majority of coding work"
  • Enforcement supplements existing Terms of Service violations

Controversy and Criticisms

The AI research community has raised significant concerns:

Invisibility Problem: Unlike transparent measures like fallback-routing, users receive no notification when RSI suppression activates, creating uncertainty about model capabilities versus artificial restrictions.

Research Interference: Academic and commercial AI research may be unknowingly compromised, affecting innovation and competitive dynamics in frontier AI development.

Trust Erosion: Silent modifications undermine confidence in model consistency and reliability for professional applications requiring predictable behavior.

Rationale

anthropic justifies RSI suppression as targeting "the actors most willing to violate" Terms of Service restrictions on developing competing models, arguing that transparent enforcement would be less effective against bad actors while silent enforcement avoids accelerating irresponsible AI development.

Alternative Approaches

Contrasts with fallback-routing, where risky queries are transparently redirected to less capable models with clear user notification, preserving trust while maintaining safety objectives.

See also

Three-Challenge Framework

page dédiée →

A fundamental framework for understanding distributed training optimization, identifying three core challenges that all scaling techniques must address: memory usage, compute efficiency, and communication overhead. This framework provides a systematic lens for evaluating and designing distributed training strategies.

The Three Core Challenges

1. Memory Usage (Hard Constraint)

"This is a hard limitation — if a training step doesn't fit in memory, training cannot proceed."

Characteristics:

  • Binary constraint: Either fits or doesn't - no partial solutions
  • Foundational requirement: Must be solved before other optimizations matter
  • Four memory components: Model weights, gradients, optimizer states, activations
  • Scaling challenge: Memory requirements grow with model size and batch size

Impact on scaling:

  • Determines maximum model size trainable on given hardware
  • Constrains batch size choices and training throughput
  • Forces architectural decisions about model and data parallelism
  • Creates hard limits that cannot be overcome through software optimization alone

2. Compute Efficiency

"We want our hardware to spend most time computing, so we need to reduce time spent on data transfers or waiting for other GPUs to perform work."

Optimization targets:

  • Maximize GPU utilization: Keep compute units busy
  • Minimize idle time: Reduce waiting periods between operations
  • Reduce data transfer overhead: Optimize memory bandwidth usage
  • Balance workload distribution: Prevent GPU starvation or overload

Efficiency metrics:

  • GPU utilization percentage
  • Time spent in computation vs. communication
  • Throughput (tokens processed per second)
  • Hardware efficiency (utilization relative to theoretical peak)

3. Communication Overhead

"We want to minimize communication overhead, as it keeps GPUs idle."

Communication optimization:

  • Bandwidth utilization: Make optimal use of available network capacity
  • Latency reduction: Minimize round-trip times for synchronization
  • Overlap strategies: Hide communication behind computation when possible
  • Topology awareness: Leverage fast intra-node vs. slower inter-node connections

Scaling considerations:

  • Communication overhead typically grows with number of GPUs
  • Network topology becomes critical at large scales
  • Bandwidth becomes bottleneck as clusters grow
  • Synchronization points create scaling barriers

Framework Trade-offs

Computation-Memory Trade-offs

Example: activation-recomputation

  • Trade: Increase computation to reduce memory usage
  • Mechanism: Recompute forward pass activations during backward pass
  • Result: 50-90% memory reduction for 15-20% computation increase

Communication-Computation Trade-offs

Example: Tensor Parallelism

  • Trade: Increase communication to distribute computation
  • Mechanism: Split model layers across multiple GPUs
  • Result: Reduced per-GPU memory but increased communication volume

Memory-Communication Trade-offs

Example: Pipeline Parallelism

  • Trade: Increase communication complexity to distribute memory requirements
  • Mechanism: Split model layers across pipeline stages
  • Result: Lower per-stage memory but complex inter-stage communication

Challenge Interaction Patterns

Hierarchical Constraint Resolution

  1. Memory first: Solve memory constraints before optimizing other dimensions
  2. Compute second: Optimize utilization within memory-feasible configurations
  3. Communication third: Minimize overhead in compute-efficient setups

Interconnected Optimization

Challenges cannot be solved in isolation:

  • Memory reduction techniques affect compute patterns
  • Communication strategies impact memory usage patterns
  • Compute optimization may require communication trade-offs

Framework Application to Techniques

Data Parallelism

  • Memory: Replicates model across GPUs (increases total memory usage)
  • Compute: Parallelizes batch processing (improves efficiency)
  • Communication: Requires gradient synchronization (overhead grows with scale)

Tensor Parallelism

  • Memory: Shards model weights across GPUs (reduces per-GPU memory)
  • Compute: Parallelizes matrix operations (can improve efficiency)
  • Communication: High communication frequency (potential bottleneck)

Pipeline Parallelism

  • Memory: Distributes layers across stages (reduces per-stage memory)
  • Compute: Sequential stage processing (can create bubbles/inefficiency)
  • Communication: Stage-to-stage transfers (moderate overhead)

Practical Framework Usage

Technique Evaluation

When evaluating distributed training approaches:

  1. Assess memory impact: Does it reduce, maintain, or increase memory requirements?
  2. Analyze compute efficiency: What is the utilization impact?
  3. Measure communication cost: How much overhead does it introduce?

Configuration Optimization

For finding optimal training configurations:

  1. **Start with memory

Token-Based Batch Sizing

page dédiée →

Method of measuring and specifying batch sizes in terms of total tokens processed rather than number of sequences, enabling consistent comparison of training configurations across different sequence lengths and making memory planning more predictable.

Core Concept

Traditional batch sizing counts sequences (samples), but sequences can vary dramatically in length, making memory usage and training metrics difficult to compare. Token-based batch sizing provides a standardized measurement that accounts for actual computational work performed.

Mathematical Relationship

The relationship between different batch size measurements:

batch_size_tokens = batch_size_samples × sequence_length
batch_size_samples = batch_size_tokens ÷ sequence_length

This makes training metrics independent of specific input sequence lengths used during training, enabling consistent comparison across different training configurations.

Industry Evolution and Standards

Historical Batch Size Growth

The ultra-scale-playbook documents significant evolution in batch size capabilities:

Llama 1 (2023):

  • Batch size: ~4M tokens
  • Training corpus: 1.4 trillion tokens
  • Demonstrates early large-scale training capabilities

DeepSeek (2024):

  • Batch size: ~60M tokens
  • Training corpus: 14 trillion tokens
  • Shows 15x increase in batch size capability with larger training datasets

Current Sweet Spot

Modern LLM pretraining typically operates in the range of 4-60 million tokens per batch. This range represents the optimal balance between:

  • Training Efficiency: Large enough batches to utilize hardware effectively
  • Memory Constraints: Small enough to fit in available GPU cluster memory
  • Convergence Quality: Avoiding batch sizes so large they harm model performance

Dynamic Batch Size Strategies

Gradual Scaling Approach

Advanced training often employs increasing batch sizes throughout training:

DeepSeek-V3/R1 Strategy:

  • Start: 3,072 sequences for early training
  • Increase: Up to 15,360 sequences for first 469B tokens
  • Maintain: 15,360 sequences for remaining training

Convergence Optimization

Batch size progression serves different training phases:

Early Training (Small Batches):

  • Rapid movement through training landscape
  • Quick exploration of parameter space
  • Noisy gradients help escape local minima

Later Training (Large Batches):

  • More accurate gradient estimates
  • Stable convergence to optimal performance
  • Reduced gradient noise for fine-tuning

Memory Planning Advantages

Predictable Memory Usage

Token-based measurement enables more accurate memory planning:

  • Sequence Length Independence: Memory estimates don't depend on variable sequence lengths
  • Consistent Comparison: Compare memory requirements across different model configurations
  • Scalability Planning: Predict memory needs for different cluster sizes

Memory Component Scaling

Different memory components scale predictably with token count:

  • Activations: Scale directly with batch size in tokens
  • Gradients: Fixed per model, independent of batch size
  • Model Parameters: Fixed per model, independent of batch size
  • Optimizer States: Fixed per model, independent of batch size

Training Efficiency Implications

Optimizer Step Reduction

Larger token-based batch sizes reduce total optimizer steps:

  • Fewer Steps: Same dataset coverage with fewer parameter updates
  • Reduced Overhead: Less time spent on optimizer computations
  • Better Hardware Utilization: More computation per synchronization point

Throughput Optimization

Token-based measurement aids throughput analysis:

  • Tokens per Second: Standard metric for training speed comparison
  • Hardware Utilization: Measure efficiency across different configurations
  • Cost Analysis: Compare training costs on consistent token basis

Sensitivity Analysis

Research shows model performance has relatively low sensitivity to exact batch size around optimal values, providing flexibility in:

  • Hardware Constraints: Adjust batch size based on available memory
  • Cost Optimization: Trade batch size for training time based on resource costs
  • Infrastructure Scaling: Adapt to different cluster configurations

Implementation Considerations

Memory Planning Process

  1. Determine Target Token Count: Based on model size and training goals
  2. Calculate Sequence Requirements: Divide by target sequence length
  3. Validate Memory Constraints: Check against available GPU memory
  4. Optimize Distribution: Plan parallelization strategy for target batch size

Dynamic Adjustment Strategies

  • Progressive Scaling: Gradually increase batch size during training
  • Hardware Adaptation: Adjust based on available cluster resources
  • Performance Monitoring: Track convergence quality with different batch sizes
  • Cost Optimization: Balance training speed against infrastructure costs

Cross-Configuration Comparison

Token-based measurement enables:

  • Model Scaling Studies: Compare training efficiency across model sizes
  • Infrastructure Evaluation: Assess different cluster configurations
  • Cost Analysis: Compare training approaches on consistent computational basis
  • Performance Benchmarking: Standardize metrics across different implementations

This standardized approach to batch size measurement, combined with understanding of dynamic scaling strategies, provides the foundation for effective training planning and resource optimization in modern LLM development.

See also

Ultra-Scale Playbook

page dédiée →

Comprehensive 240+ page methodology and knowledge base for scaling LLM training from single GPUs to thousands of coordinated GPUs, developed by hugging-face through systematic empirical research. Represents the first open-sourcing of previously proprietary distributed training knowledge held within elite industry labs.

Knowledge Democratization Mission

The Ultra-Scale Playbook addresses a critical gap in the AI training ecosystem: while foundation models are openly available, the knowledge and techniques for training them at scale remained "well kept within a handful of big industry labs." This comprehensive resource lifts the veil on distributed training methodologies that were previously proprietary.

Three-Pillar Foundation

The playbook is built on three complementary foundations:

1. Theoretical Understanding

  • Quick introductions to concepts and methods
  • High-level explanations of advantages and limitations
  • Memory breakdown analysis for Transformer models
  • Understanding of when and why memory constraints occur

2. Clear Code Implementations

  • picotron: Educational implementations in single, self-contained files for learning
  • nanotron: Production-ready codebase used at Hugging Face
  • Theory-to-code translations revealing implementation details and edge cases

3. Real Training Efficiency Benchmarks

  • Over 4,000 systematic scaling experiments
  • Up to 512 GPU cluster configurations tested
  • Infrastructure-specific optimization guidance
  • Reproducible performance measurements

Three Core Challenges Framework

All distributed training techniques address one or more of these fundamental challenges:

  1. Memory Usage: Hard constraint - if a training step doesn't fit in memory, training cannot proceed
  2. Compute Efficiency: Maximizing hardware utilization, minimizing idle time and data transfer delays
  3. Communication Overhead: Reducing inter-GPU communication that keeps devices idle, optimizing bandwidth usage

Memory as Hard Constraint

The playbook establishes memory as the primary bottleneck in LLM training, consisting of four critical components:

  • Model Weights: Parameters of the neural network
  • Gradients: Computed during backward pass
  • Optimizer States: Often the largest component (e.g., Adam momentum and variance)
  • Activations: Intermediate values needed for gradient computation

Training Step Anatomy

Detailed analysis of what happens during a single training step:

  1. Forward Pass: Activations build up as inputs pass through layers
  2. Backward Pass: Gradients computed while activations progressively cleared
  3. Optimization Step: All gradients needed, optimizer states updated

First Step Anomaly

The first training step exhibits different memory patterns due to PyTorch caching allocator preparation work. This can lead to successful first steps followed by OOM failures in subsequent steps due to optimizer state buildup.

Batch Size Evolution

Modern LLM training uses token-based batch sizes for sequence-length independence:

  • Llama 1: ~4M tokens per batch, 1.4 trillion total tokens
  • DeepSeek: ~60M tokens per batch, 14 trillion total tokens
  • Sweet Spot: 4-60 million tokens per batch for current LLM training

Educational Philosophy

Combines systematic empirical research with educational accessibility. The approach prioritizes understanding over optimization, making complex distributed training concepts accessible through clear explanations, focused implementations, and real-world benchmarking data.

Empirical Research Foundation

Built on systematic experimentation rather than purely theoretical analysis:

  • Over 4,100 distributed experiments (16k+ including test runs)
  • Systematic scanning of distributed training layouts and model sizes
  • Infrastructure-specific optimization insights
  • Reproducible benchmarking methodologies

Impact

Represents a fundamental shift in knowledge sharing for AI training, moving previously proprietary expertise into the open-source domain. Enables researchers and practitioners to understand and implement ultra-scale training without starting from scratch or reverse-engineering techniques from scattered papers.

See also


ZeRO Optimizer

page dédiée →

Zero Redundancy Optimizer (ZeRO) is an advanced memory optimization technique for distributed training that eliminates memory redundancy by partitioning optimizer states, gradients, and parameters across devices while maintaining training efficiency.

Core Problem: Memory Redundancy

Traditional Data Parallelism Issues

In standard distributed training:

  • Each GPU maintains complete copy of model parameters
  • Each GPU stores full optimizer states (often 2-3x parameter size)
  • Each GPU accumulates complete gradient set
  • Result: Massive memory redundancy across devices

Memory Components

For a model with P parameters using Adam optimizer:

  • Model parameters: P values
  • Gradients: P values
  • Optimizer states: 2P values (momentum + variance)
  • Total per GPU: 4P values × number of GPUs

ZeRO Stages

Stage 1: Optimizer State Partitioning

  • Partition: Optimizer states across devices
  • Memory reduction: 4x reduction for Adam optimizer
  • Communication: Gather required states during optimization
  • Benefit: Significant memory savings with minimal overhead

Stage 2: Gradient Partitioning

  • Partition: Gradients in addition to optimizer states
  • Memory reduction: 8x reduction total
  • Communication: All-reduce only assigned gradient partitions
  • Synchronization: Gradients distributed and synchronized efficiently

Stage 3: Parameter Partitioning

  • Partition: Model parameters across devices
  • Memory reduction: Linear with number of devices
  • Communication: Gather parameters as needed for forward/backward
  • Complexity: Most aggressive but requires careful implementation

Implementation Strategy

Dynamic Parameter Management

Stage 3 requires sophisticated parameter handling:

  1. Forward pass: Gather required parameters just before computation
  2. Computation: Execute with temporarily assembled parameters
  3. Cleanup: Discard non-local parameters to free memory
  4. Backward pass: Repeat gathering for gradient computation

Communication Optimization

  • Overlap: Hide parameter gathering with computation
  • Prefetching: Anticipate parameter needs for next layers
  • Bucketing: Group small parameters for efficient communication

Memory Efficiency Gains

Theoretical Reductions

For N devices:

  • Stage 1: Memory per device = (P + P + 2P/N) = (2P + 2P/N)
  • Stage 2: Memory per device = (P + P/N + 2P/N) = (P + 3P/N)
  • Stage 3: Memory per device = (P/N + P/N + 2P/N) = 4P/N

Practical Benefits

  • Larger models: Train models that wouldn't fit in aggregate GPU memory
  • Bigger batches: Use memory savings for increased batch sizes
  • Longer sequences: Handle extended context lengths
  • More devices: Scale to larger numbers of GPUs effectively

Communication Patterns

All-Gather Operations

  • Frequency: Parameter gathering before each layer computation
  • Size: Only required parameter subset
  • Optimization: Overlap with computation when possible

All-Reduce for Gradients

  • Stage 1 & 2: Traditional gradient synchronization
  • Stage 3: Reduced communication volume due to partitioning
  • Bucketing: Efficient handling of small gradient groups

Trade-offs and Considerations

Communication Overhead

  • Increased frequency: More communication operations per training step
  • Network sensitivity: Performance heavily dependent on interconnect bandwidth
  • Latency impact: Higher communication latency affects training speed

Implementation Complexity

  • Stage progression: Each stage adds implementation complexity
  • Memory management: Sophisticated dynamic allocation required
  • Debugging difficulty: Distributed state makes debugging challenging

Framework Integration

DeepSpeed Implementation

  • Native support: ZeRO is core feature of Microsoft's DeepSpeed
  • Automatic optimization: Framework handles communication scheduling
  • Configuration: Simple parameter selection for different stages

Other Framework Support

  • PyTorch FSDP: Similar concepts in Fully Sharded Data Parallel
  • FairScale: Facebook's implementation of sharding strategies
  • Custom implementations: Framework-agnostic manual implementation possible

Performance Optimization

Stage Selection Strategy

Choose optimal stage based on:

  • Memory pressure: How severely memory constrained
  • Network bandwidth: Available inter-device communication
  • Model size: Larger models benefit more from aggressive stages
  • Batch size requirements: Memory needs for target batch size

Hybrid Approaches

  • Selective partitioning: Partition only specific components
  • Gradient accumulation: Combine with micro-batching strategies
  • Mixed precision: Coordinate with FP16/BF16 optimizations

Advanced Optimizations

ZeRO-Offload

  • CPU offloading: Move optimizer states to CPU memory
  • Heterogeneous memory: Utilize both GPU and CPU memory hierarchies
  • Bandwidth management: Balance GPU-CPU transfer costs

ZeRO-Infinity

  • NVMe integration: Use high-speed storage for parameter swapping
  • Memory hierarchy: GPU → CPU → NVMe memory management
  • Extremely large models: Train models larger than total system memory

See also