~/wiki

Batch Size Optimization

Mis à jour le 2025-12-31Confiance : high
batch-sizetraining-optimizationconvergencethroughputtoken-based-batchinggradient-noisebatch-size-evolutionllm-trainingmemory-constraintsoptimizer-stepsbatch-size-sensitivitysequence-length-independencebatch-size-tokensbatch-size-samples

The process of selecting optimal batch sizes for neural network training, balancing convergence properties, training efficiency, and memory constraints. In LLM training, batch size optimization has evolved significantly with models trained on increasingly large batch sizes measured in tokens.

Batch Size Impact on Training

Convergence Properties

Small Batch Sizes:

  • Useful early in training for quick movement through training landscape
  • Maintain gradient noise which can help reach optimal learning points
  • Keep gradients noisy later in training, potentially preventing optimal convergence
  • Enable faster iteration through training data

Large Batch Sizes:

  • Provide very accurate gradient estimations
  • Reduce gradient noise for stable convergence
  • May make less efficient use of each training token
  • Can lead to slower convergence despite better gradient estimates

Training Efficiency Impact

Batch size affects total training time through optimizer step frequency:

  • Small batches: Require more optimizer steps for same amount of data
  • Large batches: Fewer optimizer steps but potentially less efficient token utilization
  • Optimizer step cost: Each step is computationally expensive independent of batch size

Token-Based Reporting Convention

Industry Standard

LLM pretraining community reports batch sizes in tokens rather than samples to maintain independence from sequence length variations.

Conversion Formula

batch_size_tokens = batch_size_samples × sequence_length

This allows training comparisons across different sequence length configurations.

Historical Batch Size Evolution

Major Model Progression

  • Llama 1: ~4M tokens per batch, 1.4T total tokens
  • DeepSeek: ~60M tokens per batch, 14T total tokens
  • DeepSeek-V3/R1: Gradual scaling from 3,072 to 15,360 sequences over first 469B tokens

Modern Sweet Spot

Current optimal range for LLM training: 4-60 million tokens per batch

Dynamic Batch Size Strategies

Gradual Batch Size Increase

Example from DeepSeek-V3/R1 training:

  1. Initial phase: 3,072 input sequences
  2. Scaling phase: Gradually increase to 15,360 sequences over first 469B tokens
  3. Steady phase: Maintain 15,360 sequences for remaining training

This approach balances early training exploration with later training stability.

Batch Size Sensitivity Analysis

Low Sensitivity Around Optimum

"The sensitivity of final model performance to the exact batch size value is usually rather low around the optimal batch size"

Practical Implications

  • Batch size can be adjusted within reasonable ranges without major performance impact
  • Allows optimization for hardware constraints and memory limitations
  • Enables batch size tuning for infrastructure-specific requirements

Memory Constraints and Batch Size

Out-of-Memory Challenge

Primary limitation when scaling to large batch sizes: GPU memory capacity

Memory-Batch Size Relationship

Larger batch sizes increase memory requirements through:

  • Activation storage: Scales linearly with batch size
  • Gradient accumulation: More samples require more intermediate storage
  • Attention computation: Quadratic scaling with sequence length

Optimization Strategies

Gradient Accumulation

When target batch size exceeds memory capacity:

  1. Split large batch into smaller micro-batches
  2. Accumulate gradients across micro-batches
  3. Update parameters after processing full logical batch
  4. Achieves large effective batch size within memory constraints

Memory-Efficient Batch Sizing

  • Profile memory usage at different batch sizes
  • Find maximum batch size that fits in available memory
  • Use gradient accumulation to reach target effective batch size
  • Balance memory usage with training throughput

Practical Batch Size Selection

Three-Step Process

  1. Determine target batch size based on model size and training objectives
  2. Assess memory constraints to find maximum feasible batch size
  3. Apply gradient accumulation if needed to bridge the gap

Infrastructure Considerations

  • Available GPU memory per device
  • Number of GPUs in training setup
  • Model parallelization strategy
  • Sequence length requirements

See Also