LLM Scaling Techniques
Methods and strategies for training increasingly large language models, encompassing both model architecture scaling and training infrastructure scaling. Critical for developing state-of-the-art AI systems that require coordination across hundreds to thousands of GPUs.
Batch Size Scaling Evolution
The LLM training community has seen dramatic increases in batch sizes over time, reflecting improved distributed training capabilities:
- Llama 1: ~4M tokens per batch, 1.4 trillion total training tokens
- DeepSeek: ~60M tokens per batch, 14 trillion total training tokens
- DeepSeek-V3/R1: Dynamic scaling from 3,072 to 15,360 input sequences during initial 469B tokens, then maintained at 15,360
Token-Based Measurement
Modern LLM training reports batch sizes in tokens rather than samples to maintain independence from sequence length:
batch_size_tokens = batch_size_samples × sequence_length
This standardization enables consistent comparison across different model architectures and training configurations.
Three-Challenge Framework
All scaling techniques address one or more of three fundamental challenges identified in the ultra-scale-playbook:
1. Memory Usage (Hard Constraint)
- Nature: If training step doesn't fit in memory, training cannot proceed
- Components: Model weights, gradients, optimizer states, activations
- Solutions: memory-optimization, activation-recomputation, gradient accumulation
2. Compute Efficiency
- Goal: Maximize hardware utilization by reducing idle time
- Challenges: Data transfer overhead, GPU synchronization delays
- Solutions: Kernel fusion, mixed precision, optimized data pipelines
3. Communication Overhead
- Impact: Inter-GPU communication keeps hardware idle
- Strategy: Optimize intra-node (fast) vs inter-node (slower) bandwidth usage
- Approach: Overlap communication with computation whenever possible
Five-Dimensional Parallelism
Modern LLM training requires coordination across multiple parallelization dimensions:
- Data Parallelism: Distribute training data across GPUs
- Tensor Parallelism: Split model tensors across devices
- Pipeline Parallelism: Distribute model layers across pipeline stages
- Context Parallelism: Handle sequences longer than single-device memory
- Expert Parallelism: Distribute mixture-of-experts layers
Trade-off Optimization
Scaling techniques often involve trading one resource against another:
- Computation vs Memory: activation-recomputation trades 15-20% additional compute for 50-90% memory reduction
- Communication vs Parallelism: Higher parallelism increases communication overhead
- Batch Size vs Convergence: Larger batches improve hardware utilization but may affect convergence properties
Empirical Research Foundation
The ultra-scale-playbook provides data-driven optimization through over 4,000 scaling experiments:
- Up to 512 GPUs tested
- Systematic measurement of throughput and GPU utilization
- Normalized performance metrics across different model sizes
- Reproducible benchmarking methodology
Batch Size Sensitivity and Optimization
Convergence Considerations
Batch size affects model convergence in complex ways:
- Small batches: Quick initial progress but noisy gradients prevent optimal final performance
- Large batches: Accurate gradients but potentially slower convergence and compute waste
- Sweet spot: Typically 4-60 million tokens per batch for recent LLMs
Dynamic Scaling Strategies
Advanced training runs employ dynamic batch size scaling:
- Start with smaller batches for rapid initial convergence
- Gradually increase batch size as training progresses
- Balance convergence speed with computational efficiency
Memory-First Design Philosophy
Modern scaling techniques prioritize memory optimization since memory represents a hard constraint unlike compute or communication which can be optimized without preventing training:
- Memory prediction tools for capacity planning
- Systematic memory component analysis
- training-step-anatomy understanding for optimization
- Empirical memory profiling for validation
See also
- ultra-scale-playbook
- gpu-cluster-training
- memory-optimization
- Five-Dimensional Parallelism
- training-step-anatomy
- distributed-training