---
title: Attention Mechanisms
category: concepts
created: 2026-06-11
updated: 2026-12-20
tags: [attention, self-attention, cross-attention, transformer, query-key-value, attention-weights, sequence-modeling, neural-networks, mathematical-notation, attention-matrix]
sources: [raw/feeds/2026-06-11-the-transformer-family-version-2-0.md]
confidence: high
---
# Attention Mechanisms
Core computational mechanism in modern neural networks that allows models to selectively focus on different parts of the input when processing sequences. The foundation of the [transformer-architecture](/concepts/transformer-architecture) and essential for understanding how large language models process information.
## Mathematical Foundation
### Query-Key-Value Framework
Based on lilian-weng's comprehensive notation:
- **Query matrix**: $\mathbf{Q} = \mathbf{X}\mathbf{W}^q \in \mathbb{R}^{L \times d_k}$
- **Key matrix**: $\mathbf{K} = \mathbf{X}\mathbf{W}^k \in \mathbb{R}^{L \times d_k}$
- **Value matrix**: $\mathbf{V} = \mathbf{X}\mathbf{W}^v \in \mathbb{R}^{L \times d_v}$
Where $\mathbf{X} \in \mathbb{R}^{L \times d}$ is the input sequence and weight matrices $\mathbf{W}^q, \mathbf{W}^k \in \mathbb{R}^{d \times d_k}$, $\mathbf{W}^v \in \mathbb{R}^{d \times d_v}$ transform inputs.
### Attention Computation
The self-attention matrix between input sequence and itself:
$$\mathbf{A} = \text{softmax}\left(\frac{\mathbf{Q}\mathbf{K}^\top}{\sqrt{d_k}}\right) \in \mathbb{R}^{L \times L}$$
Each element $a_{ij}$ represents the attention score between query $\mathbf{q}_i$ and key $\mathbf{k}_j$.
## Types of Attention
### Self-Attention
- Queries, keys, and values all derived from the same input sequence
- Enables modeling relationships within a single sequence
- Foundation of transformer encoder and decoder layers
### Cross-Attention
- Queries from one sequence, keys/values from another
- Used in encoder-decoder architectures for translation tasks
- Enables alignment between source and target sequences
## Key Properties
- **Parallel Processing**: All positions attended to simultaneously
- **Position-Agnostic**: Requires [positional-encoding](/concepts/positional-encoding) for sequence order
- **Flexible Context**: Each position can attend to any other position
- **Weighted Representation**: Output is weighted combination of values
## See also
- [multi-head-attention](/concepts/multi-head-attention)
- [transformer-architecture](/concepts/transformer-architecture)
- [positional-encoding](/concepts/positional-encoding)
- lilian-weng