~/wiki

Attention Matrix Computation

Mis à jour le 2025-01-04Confiance : high
attention-matrixattention-computationself-attentionquery-key-valuesoftmaxattention-scorestransformer-architecturemathematical-computation

Core computational process in transformer-architecture that produces attention weights determining how much focus each position pays to other positions in the sequence. The foundation of Self-Attention mechanisms.

Mathematical Formulation

Attention Matrix

A ∈ R^{L×L} = softmax(QK^T / √d_k)

Where:

  • A: Self-attention matrix between input sequence of length L and itself
  • Q: Query embeddings Q = XW^q
  • K: Key embeddings K = XW^k
  • L: Sequence length
  • d_k: Key dimension

Scalar Attention Score

a_ij ∈ A

Represents the attention score between query q_i and key k_j, determining how much position i attends to position j.

Computation Steps

  1. Query-Key Interaction: Compute dot products QK^T
  2. Scaling: Divide by √d_k to prevent softmax saturation
  3. Normalization: Apply softmax to get probability distribution
  4. Output: Weighted combination with values AV

Key Properties

  • Symmetrical: Can attend to any position in the sequence
  • Normalized: Each row sums to 1 after softmax
  • Differentiable: Enables gradient-based training
  • Parallel: All positions computed simultaneously

See also