Batch Size Optimization
The process of selecting optimal batch sizes for neural network training, balancing convergence properties, training efficiency, and memory constraints. In LLM training, batch size optimization has evolved significantly with models trained on increasingly large batch sizes measured in tokens.
Batch Size Impact on Training
Convergence Properties
Small Batch Sizes:
- Useful early in training for quick movement through training landscape
- Maintain gradient noise which can help reach optimal learning points
- Keep gradients noisy later in training, potentially preventing optimal convergence
- Enable faster iteration through training data
Large Batch Sizes:
- Provide very accurate gradient estimations
- Reduce gradient noise for stable convergence
- May make less efficient use of each training token
- Can lead to slower convergence despite better gradient estimates
Training Efficiency Impact
Batch size affects total training time through optimizer step frequency:
- Small batches: Require more optimizer steps for same amount of data
- Large batches: Fewer optimizer steps but potentially less efficient token utilization
- Optimizer step cost: Each step is computationally expensive independent of batch size
Token-Based Reporting Convention
Industry Standard
LLM pretraining community reports batch sizes in tokens rather than samples to maintain independence from sequence length variations.
Conversion Formula
batch_size_tokens = batch_size_samples × sequence_length
This allows training comparisons across different sequence length configurations.
Historical Batch Size Evolution
Major Model Progression
- Llama 1: ~4M tokens per batch, 1.4T total tokens
- DeepSeek: ~60M tokens per batch, 14T total tokens
- DeepSeek-V3/R1: Gradual scaling from 3,072 to 15,360 sequences over first 469B tokens
Modern Sweet Spot
Current optimal range for LLM training: 4-60 million tokens per batch
Dynamic Batch Size Strategies
Gradual Batch Size Increase
Example from DeepSeek-V3/R1 training:
- Initial phase: 3,072 input sequences
- Scaling phase: Gradually increase to 15,360 sequences over first 469B tokens
- Steady phase: Maintain 15,360 sequences for remaining training
This approach balances early training exploration with later training stability.
Batch Size Sensitivity Analysis
Low Sensitivity Around Optimum
"The sensitivity of final model performance to the exact batch size value is usually rather low around the optimal batch size"
Practical Implications
- Batch size can be adjusted within reasonable ranges without major performance impact
- Allows optimization for hardware constraints and memory limitations
- Enables batch size tuning for infrastructure-specific requirements
Memory Constraints and Batch Size
Out-of-Memory Challenge
Primary limitation when scaling to large batch sizes: GPU memory capacity
Memory-Batch Size Relationship
Larger batch sizes increase memory requirements through:
- Activation storage: Scales linearly with batch size
- Gradient accumulation: More samples require more intermediate storage
- Attention computation: Quadratic scaling with sequence length
Optimization Strategies
Gradient Accumulation
When target batch size exceeds memory capacity:
- Split large batch into smaller micro-batches
- Accumulate gradients across micro-batches
- Update parameters after processing full logical batch
- Achieves large effective batch size within memory constraints
Memory-Efficient Batch Sizing
- Profile memory usage at different batch sizes
- Find maximum batch size that fits in available memory
- Use gradient accumulation to reach target effective batch size
- Balance memory usage with training throughput
Practical Batch Size Selection
Three-Step Process
- Determine target batch size based on model size and training objectives
- Assess memory constraints to find maximum feasible batch size
- Apply gradient accumulation if needed to bridge the gap
Infrastructure Considerations
- Available GPU memory per device
- Number of GPUs in training setup
- Model parallelization strategy
- Sequence length requirements
See Also
- memory-optimization
- gradient-accumulation
- training-step-anatomy
- ultra-scale-playbook
- Token-Based Training Metrics