~/wiki

Mathematical Notation for Transformers

Mis à jour le 2025-01-04Confiance : high
mathematical-notationtransformer-architecturestandardizationtechnical-communicationmatrix-dimensionsattention-computationlilian-wengquery-key-valuepositional-encodingcomprehensive-notation-systemtransformer-family-v2version-2-0model-parametersweight-matrices

Standardized mathematical notation system for transformer-architecture components, essential for precise technical communication and implementation. lilian-weng's comprehensive Version 2.0 notation provides the mathematical rigor needed to understand transformer variants and their architectural modifications.

Core Notation System

Model Dimensions

  • $d$: Model size/hidden state dimension/positional encoding size
  • $h$: Number of heads in multi-head attention layer
  • $L$: Segment length of input sequence
  • $N$: Total number of attention layers (excluding MoE)

Input and Matrices

  • $\mathbf{X} \in \mathbb{R}^{L \times d}$: Input sequence with embedding vectors
  • $\mathbf{P} \in \mathbb{R}^{L \times d}$: positional-encoding matrix

Weight Matrices

  • $\mathbf{W}^q \in \mathbb{R}^{d \times d_k}$: Query weight matrix
  • $\mathbf{W}^k \in \mathbb{R}^{d \times d_k}$: Key weight matrix
  • $\mathbf{W}^v \in \mathbb{R}^{d \times d_v}$: Value weight matrix
  • $\mathbf{W}^o \in \mathbb{R}^{d_v \times d}$: Output weight matrix

Multi-Head Components

  • $\mathbf{W}^k_i, \mathbf{W}^q_i \in \mathbb{R}^{d \times d_k/h}$: Per-head key and query weights
  • $\mathbf{W}^v_i \in \mathbb{R}^{d \times d_v/h}$: Per-head value weights

Attention Computation

  • $\mathbf{Q} = \mathbf{X}\mathbf{W}^q \in \mathbb{R}^{L \times d_k}$: Query embeddings
  • $\mathbf{K} = \mathbf{X}\mathbf{W}^k \in \mathbb{R}^{L \times d_k}$: Key embeddings
  • $\mathbf{V} = \mathbf{X}\mathbf{W}^v \in \mathbb{R}^{L \times d_v}$: Value embeddings
  • $\mathbf{A} \in \mathbb{R}^{L \times L}$: Self-attention matrix
  • $a_{ij} \in \mathbf{A}$: Scalar attention score between query $i$ and key $j$

Position and Attention Sets

  • $\mathbf{q}_i, \mathbf{k}_i \in \mathbb{R}^{d_k}, \mathbf{v}_i \in \mathbb{R}^{d_v}$: Row vectors in Q, K, V matrices
  • $S_i$: Collection of key positions for query $i$ to attend to

Standardization Benefits

  1. Precision: Eliminates ambiguity in architectural descriptions
  2. Consistency: Unified symbols across transformer literature
  3. Implementation: Direct mapping to code implementations
  4. Research Communication: Clear technical discourse foundation

See also