Token-Based Batch Sizing
Method of measuring and specifying batch sizes in terms of total tokens processed rather than number of sequences, enabling consistent comparison of training configurations across different sequence lengths and making memory planning more predictable.
Core Concept
Traditional batch sizing counts sequences (samples), but sequences can vary dramatically in length, making memory usage and training metrics difficult to compare. Token-based batch sizing provides a standardized measurement that accounts for actual computational work performed.
Mathematical Relationship
The relationship between different batch size measurements:
batch_size_tokens = batch_size_samples × sequence_length
batch_size_samples = batch_size_tokens ÷ sequence_length
This makes training metrics independent of specific input sequence lengths used during training, enabling consistent comparison across different training configurations.
Industry Evolution and Standards
Historical Batch Size Growth
The ultra-scale-playbook documents significant evolution in batch size capabilities:
Llama 1 (2023):
- Batch size: ~4M tokens
- Training corpus: 1.4 trillion tokens
- Demonstrates early large-scale training capabilities
DeepSeek (2024):
- Batch size: ~60M tokens
- Training corpus: 14 trillion tokens
- Shows 15x increase in batch size capability with larger training datasets
Current Sweet Spot
Modern LLM pretraining typically operates in the range of 4-60 million tokens per batch. This range represents the optimal balance between:
- Training Efficiency: Large enough batches to utilize hardware effectively
- Memory Constraints: Small enough to fit in available GPU cluster memory
- Convergence Quality: Avoiding batch sizes so large they harm model performance
Dynamic Batch Size Strategies
Gradual Scaling Approach
Advanced training often employs increasing batch sizes throughout training:
DeepSeek-V3/R1 Strategy:
- Start: 3,072 sequences for early training
- Increase: Up to 15,360 sequences for first 469B tokens
- Maintain: 15,360 sequences for remaining training
Convergence Optimization
Batch size progression serves different training phases:
Early Training (Small Batches):
- Rapid movement through training landscape
- Quick exploration of parameter space
- Noisy gradients help escape local minima
Later Training (Large Batches):
- More accurate gradient estimates
- Stable convergence to optimal performance
- Reduced gradient noise for fine-tuning
Memory Planning Advantages
Predictable Memory Usage
Token-based measurement enables more accurate memory planning:
- Sequence Length Independence: Memory estimates don't depend on variable sequence lengths
- Consistent Comparison: Compare memory requirements across different model configurations
- Scalability Planning: Predict memory needs for different cluster sizes
Memory Component Scaling
Different memory components scale predictably with token count:
- Activations: Scale directly with batch size in tokens
- Gradients: Fixed per model, independent of batch size
- Model Parameters: Fixed per model, independent of batch size
- Optimizer States: Fixed per model, independent of batch size
Training Efficiency Implications
Optimizer Step Reduction
Larger token-based batch sizes reduce total optimizer steps:
- Fewer Steps: Same dataset coverage with fewer parameter updates
- Reduced Overhead: Less time spent on optimizer computations
- Better Hardware Utilization: More computation per synchronization point
Throughput Optimization
Token-based measurement aids throughput analysis:
- Tokens per Second: Standard metric for training speed comparison
- Hardware Utilization: Measure efficiency across different configurations
- Cost Analysis: Compare training costs on consistent token basis
Sensitivity Analysis
Research shows model performance has relatively low sensitivity to exact batch size around optimal values, providing flexibility in:
- Hardware Constraints: Adjust batch size based on available memory
- Cost Optimization: Trade batch size for training time based on resource costs
- Infrastructure Scaling: Adapt to different cluster configurations
Implementation Considerations
Memory Planning Process
- Determine Target Token Count: Based on model size and training goals
- Calculate Sequence Requirements: Divide by target sequence length
- Validate Memory Constraints: Check against available GPU memory
- Optimize Distribution: Plan parallelization strategy for target batch size
Dynamic Adjustment Strategies
- Progressive Scaling: Gradually increase batch size during training
- Hardware Adaptation: Adjust based on available cluster resources
- Performance Monitoring: Track convergence quality with different batch sizes
- Cost Optimization: Balance training speed against infrastructure costs
Cross-Configuration Comparison
Token-based measurement enables:
- Model Scaling Studies: Compare training efficiency across model sizes
- Infrastructure Evaluation: Assess different cluster configurations
- Cost Analysis: Compare training approaches on consistent computational basis
- Performance Benchmarking: Standardize metrics across different implementations
This standardized approach to batch size measurement, combined with understanding of dynamic scaling strategies, provides the foundation for effective training planning and resource optimization in modern LLM development.
See also
- ultra-scale-playbook - Comprehensive methodology including batch sizing strategies
- memory-optimization - Techniques for managing memory constraints with different batch sizes
- training-step-anatomy - How batch size affects memory patterns during training
- gpu-cluster-training - Hardware considerations for large batch training
- llm-scaling-techniques - Broader context of scaling approaches in language model training