~/wiki

Memory Bandwidth Bottleneck

Mis à jour le 2025-01-03Confiance : high
memory-bandwidthinference-optimizationhardware-constraintsgpu-optimizationperformance-bottleneckmemory-accesscomputational-efficiencytransformer-modelsautoregressive-generationsequential-processinglilian-wengpope-et-al

A fundamental performance constraint in large transformer model inference where memory access speed becomes the limiting factor rather than raw computational capability. This bottleneck occurs when the rate at which data can be transferred between memory and processing units is slower than the processor's ability to consume that data, creating a critical constraint for large model deployment.

Technical Foundation

Identified by pope-et-al-2022 and systematically analyzed by lilian-weng, this bottleneck represents one of two primary factors (along with autoregressive-generation) that make large transformer inference challenging beyond just model size considerations.

Why It Occurs

Model Size vs. Memory Speed

Large transformer models require billions of parameters to be loaded from memory, but memory bandwidth has not scaled at the same rate as model size growth. The sheer volume of data that must be transferred creates a fundamental constraint.

Sequential Access Patterns

autoregressive-generation requires sequential processing where each token generation depends on all previous tokens, creating memory access patterns that cannot be easily parallelized or cached effectively.

Hardware Limitations

Even high-end GPUs with substantial computational power are constrained by the rate at which they can access model parameters from memory, making memory bandwidth rather than FLOPS the limiting factor.

Impact on Inference

Latency Implications

Memory bandwidth constraints directly translate to increased inference latency, as the model must wait for parameters to be loaded before computation can proceed.

Throughput Limitations

Batch processing efficiency is reduced when memory access becomes the bottleneck, limiting the number of requests that can be processed simultaneously.

Resource Utilization

Computational units may remain underutilized while waiting for data, reducing overall system efficiency and increasing cost per inference.

Optimization Strategies

Model Compression

Memory Optimization

  • Parameter caching: Strategic loading and retention of frequently accessed parameters
  • Memory layout optimization: Improving data locality and access patterns
  • Mixed precision: Using different precisions to reduce bandwidth requirements

Hardware Solutions

  • High-bandwidth memory: Specialized memory architectures with increased bandwidth
  • On-chip caching: Keeping frequently accessed parameters closer to compute units
  • Memory hierarchy optimization: Leveraging different memory tiers effectively

See also