~/wiki

Autoregressive Generation

Mis à jour le 2025-01-03Confiance : high
autoregressive-generationsequential-processingtransformer-modelstoken-generationinference-constraintsnext-token-predictionlanguage-modelinglilian-wengpope-et-alparallelism-limitationsmemory-bandwidth

The fundamental approach used by most large language models where text is generated sequentially, one token at a time, with each new token conditioned on all previously generated tokens. This creates a natural language generation process but introduces significant computational constraints that limit inference-optimization strategies.

Technical Foundation

Identified by pope-et-al-2022 and analyzed by lilian-weng as one of two primary factors (along with memory-bandwidth-bottleneck) that make large transformer inference challenging beyond just model size considerations.

How It Works

Sequential Dependency

Each token generation step requires:

  1. Processing all previous tokens in the sequence
  2. Computing attention weights across the entire context
  3. Generating probability distribution over vocabulary
  4. Sampling or selecting the next token

Context Accumulation

As sequences grow longer, the computational cost increases because:

  • Attention computation scales quadratically with sequence length
  • Each step requires processing the expanding context
  • Memory requirements grow with sequence length

Inference Constraints

Parallelization Limitations

Unlike training where multiple tokens can be processed simultaneously, autoregressive inference cannot parallelize token generation within a single sequence because each token depends on all previous tokens.

Memory Access Patterns

Sequential generation creates challenging memory access patterns:

  • Key-value caches must be maintained and accessed for each step
  • Memory bandwidth becomes constrained by sequential access requirements
  • Cache management becomes critical for longer sequences

Latency Accumulation

Total inference latency is the sum of individual token generation times:

  • Each token adds to total response time
  • Longer sequences result in proportionally longer latency
  • Interactive applications are particularly sensitive to this accumulation

Optimization Strategies

Key-Value Caching

  • Store computed attention keys and values to avoid recomputation
  • Trade memory for computational efficiency
  • Critical for maintaining reasonable inference speeds

Speculative Decoding

  • Generate multiple potential tokens in parallel
  • Verify correctness against the main model
  • Can provide speedups when speculation succeeds

Parallel Sampling

  • Generate multiple independent sequences simultaneously
  • Leverage batch processing for throughput optimization
  • Doesn't solve single-sequence latency but improves overall efficiency

Context Management

  • Sliding window attention to limit context growth
  • Context compression techniques
  • Strategic context truncation strategies

Impact on System Design

Hardware Requirements

Autoregressive constraints influence hardware design priorities:

  • Memory bandwidth becomes more critical than raw compute
  • Cache hierarchy optimization is essential
  • Sequential processing limits parallelization benefits

Software Architecture

Systems must be designed around sequential constraints:

  • Stateful processing with context management
  • Efficient key-value cache implementation
  • Memory optimization for long sequences

See also