~/wiki

Token-Based Batch Sizing

Confiance : high
batch-sizingtoken-countingtraining-optimizationsequence-lengthllm-trainingmemory-managementindustry-standardsscaling-comparisonultra-scale-playbookbatch-size-evolutiondistributed-trainingmemory-planningthroughput-optimizationbatch-size-tokenssequence-length-independencehistorical-evolutiondynamic-batch-sizingoptimizer-stepsconvergence-optimizationmemory-predictabilitytraining-metrics

Method of measuring and specifying batch sizes in terms of total tokens processed rather than number of sequences, enabling consistent comparison of training configurations across different sequence lengths and making memory planning more predictable.

Core Concept

Traditional batch sizing counts sequences (samples), but sequences can vary dramatically in length, making memory usage and training metrics difficult to compare. Token-based batch sizing provides a standardized measurement that accounts for actual computational work performed.

Mathematical Relationship

The relationship between different batch size measurements:

batch_size_tokens = batch_size_samples × sequence_length
batch_size_samples = batch_size_tokens ÷ sequence_length

This makes training metrics independent of specific input sequence lengths used during training, enabling consistent comparison across different training configurations.

Industry Evolution and Standards

Historical Batch Size Growth

The ultra-scale-playbook documents significant evolution in batch size capabilities:

Llama 1 (2023):

  • Batch size: ~4M tokens
  • Training corpus: 1.4 trillion tokens
  • Demonstrates early large-scale training capabilities

DeepSeek (2024):

  • Batch size: ~60M tokens
  • Training corpus: 14 trillion tokens
  • Shows 15x increase in batch size capability with larger training datasets

Current Sweet Spot

Modern LLM pretraining typically operates in the range of 4-60 million tokens per batch. This range represents the optimal balance between:

  • Training Efficiency: Large enough batches to utilize hardware effectively
  • Memory Constraints: Small enough to fit in available GPU cluster memory
  • Convergence Quality: Avoiding batch sizes so large they harm model performance

Dynamic Batch Size Strategies

Gradual Scaling Approach

Advanced training often employs increasing batch sizes throughout training:

DeepSeek-V3/R1 Strategy:

  • Start: 3,072 sequences for early training
  • Increase: Up to 15,360 sequences for first 469B tokens
  • Maintain: 15,360 sequences for remaining training

Convergence Optimization

Batch size progression serves different training phases:

Early Training (Small Batches):

  • Rapid movement through training landscape
  • Quick exploration of parameter space
  • Noisy gradients help escape local minima

Later Training (Large Batches):

  • More accurate gradient estimates
  • Stable convergence to optimal performance
  • Reduced gradient noise for fine-tuning

Memory Planning Advantages

Predictable Memory Usage

Token-based measurement enables more accurate memory planning:

  • Sequence Length Independence: Memory estimates don't depend on variable sequence lengths
  • Consistent Comparison: Compare memory requirements across different model configurations
  • Scalability Planning: Predict memory needs for different cluster sizes

Memory Component Scaling

Different memory components scale predictably with token count:

  • Activations: Scale directly with batch size in tokens
  • Gradients: Fixed per model, independent of batch size
  • Model Parameters: Fixed per model, independent of batch size
  • Optimizer States: Fixed per model, independent of batch size

Training Efficiency Implications

Optimizer Step Reduction

Larger token-based batch sizes reduce total optimizer steps:

  • Fewer Steps: Same dataset coverage with fewer parameter updates
  • Reduced Overhead: Less time spent on optimizer computations
  • Better Hardware Utilization: More computation per synchronization point

Throughput Optimization

Token-based measurement aids throughput analysis:

  • Tokens per Second: Standard metric for training speed comparison
  • Hardware Utilization: Measure efficiency across different configurations
  • Cost Analysis: Compare training costs on consistent token basis

Sensitivity Analysis

Research shows model performance has relatively low sensitivity to exact batch size around optimal values, providing flexibility in:

  • Hardware Constraints: Adjust batch size based on available memory
  • Cost Optimization: Trade batch size for training time based on resource costs
  • Infrastructure Scaling: Adapt to different cluster configurations

Implementation Considerations

Memory Planning Process

  1. Determine Target Token Count: Based on model size and training goals
  2. Calculate Sequence Requirements: Divide by target sequence length
  3. Validate Memory Constraints: Check against available GPU memory
  4. Optimize Distribution: Plan parallelization strategy for target batch size

Dynamic Adjustment Strategies

  • Progressive Scaling: Gradually increase batch size during training
  • Hardware Adaptation: Adjust based on available cluster resources
  • Performance Monitoring: Track convergence quality with different batch sizes
  • Cost Optimization: Balance training speed against infrastructure costs

Cross-Configuration Comparison

Token-based measurement enables:

  • Model Scaling Studies: Compare training efficiency across model sizes
  • Infrastructure Evaluation: Assess different cluster configurations
  • Cost Analysis: Compare training approaches on consistent computational basis
  • Performance Benchmarking: Standardize metrics across different implementations

This standardized approach to batch size measurement, combined with understanding of dynamic scaling strategies, provides the foundation for effective training planning and resource optimization in modern LLM development.

See also