Attention Matrix Computation
Mis à jour le 2025-01-04Confiance : high
attention-matrixattention-computationself-attentionquery-key-valuesoftmaxattention-scorestransformer-architecturemathematical-computation
Core computational process in transformer-architecture that produces attention weights determining how much focus each position pays to other positions in the sequence. The foundation of Self-Attention mechanisms.
Mathematical Formulation
Attention Matrix
A ∈ R^{L×L} = softmax(QK^T / √d_k)
Where:
- A: Self-attention matrix between input sequence of length L and itself
- Q: Query embeddings
Q = XW^q - K: Key embeddings
K = XW^k - L: Sequence length
- d_k: Key dimension
Scalar Attention Score
a_ij ∈ A
Represents the attention score between query q_i and key k_j, determining how much position i attends to position j.
Computation Steps
- Query-Key Interaction: Compute dot products
QK^T - Scaling: Divide by
√d_kto prevent softmax saturation - Normalization: Apply softmax to get probability distribution
- Output: Weighted combination with values
AV
Key Properties
- Symmetrical: Can attend to any position in the sequence
- Normalized: Each row sums to 1 after softmax
- Differentiable: Enables gradient-based training
- Parallel: All positions computed simultaneously
See also
- multi-head-attention
- attention-mechanisms
- Self-Attention
- Query-Key-Value