Mathematical Notation for Transformers
Mis à jour le 2025-01-04Confiance : high
mathematical-notationtransformer-architecturestandardizationtechnical-communicationmatrix-dimensionsattention-computationlilian-wengquery-key-valuepositional-encodingcomprehensive-notation-systemtransformer-family-v2version-2-0model-parametersweight-matrices
Standardized mathematical notation system for transformer-architecture components, essential for precise technical communication and implementation. lilian-weng's comprehensive Version 2.0 notation provides the mathematical rigor needed to understand transformer variants and their architectural modifications.
Core Notation System
Model Dimensions
- $d$: Model size/hidden state dimension/positional encoding size
- $h$: Number of heads in multi-head attention layer
- $L$: Segment length of input sequence
- $N$: Total number of attention layers (excluding MoE)
Input and Matrices
- $\mathbf{X} \in \mathbb{R}^{L \times d}$: Input sequence with embedding vectors
- $\mathbf{P} \in \mathbb{R}^{L \times d}$: positional-encoding matrix
Weight Matrices
- $\mathbf{W}^q \in \mathbb{R}^{d \times d_k}$: Query weight matrix
- $\mathbf{W}^k \in \mathbb{R}^{d \times d_k}$: Key weight matrix
- $\mathbf{W}^v \in \mathbb{R}^{d \times d_v}$: Value weight matrix
- $\mathbf{W}^o \in \mathbb{R}^{d_v \times d}$: Output weight matrix
Multi-Head Components
- $\mathbf{W}^k_i, \mathbf{W}^q_i \in \mathbb{R}^{d \times d_k/h}$: Per-head key and query weights
- $\mathbf{W}^v_i \in \mathbb{R}^{d \times d_v/h}$: Per-head value weights
Attention Computation
- $\mathbf{Q} = \mathbf{X}\mathbf{W}^q \in \mathbb{R}^{L \times d_k}$: Query embeddings
- $\mathbf{K} = \mathbf{X}\mathbf{W}^k \in \mathbb{R}^{L \times d_k}$: Key embeddings
- $\mathbf{V} = \mathbf{X}\mathbf{W}^v \in \mathbb{R}^{L \times d_v}$: Value embeddings
- $\mathbf{A} \in \mathbb{R}^{L \times L}$: Self-attention matrix
- $a_{ij} \in \mathbf{A}$: Scalar attention score between query $i$ and key $j$
Position and Attention Sets
- $\mathbf{q}_i, \mathbf{k}_i \in \mathbb{R}^{d_k}, \mathbf{v}_i \in \mathbb{R}^{d_v}$: Row vectors in Q, K, V matrices
- $S_i$: Collection of key positions for query $i$ to attend to
Standardization Benefits
- Precision: Eliminates ambiguity in architectural descriptions
- Consistency: Unified symbols across transformer literature
- Implementation: Direct mapping to code implementations
- Research Communication: Clear technical discourse foundation