~/wiki

multi head attention

---
title: Multi-Head Attention
category: concepts
created: 2026-06-11
updated: 2026-12-20
tags: [multi-head-attention, transformer, attention-heads, parallel-attention, representation-subspaces, query-key-value, mathematical-notation, attention-concatenation]
sources: [raw/feeds/2026-06-11-the-transformer-family-version-2-0.md]
confidence: high
---

# Multi-Head Attention

Parallel attention mechanism in [transformer-architecture](/concepts/transformer-architecture) that runs multiple independent attention functions simultaneously, allowing the model to focus on different types of information and relationships within the input sequence.

## Mathematical Framework

### Head-Specific Parameters
Following lilian-weng's notation, each attention head $i$ has its own weight matrices:
- **Query weights**: $\mathbf{W}^q_i \in \mathbb{R}^{d \times d_k/h}$
- **Key weights**: $\mathbf{W}^k_i \in \mathbb{R}^{d \times d_k/h}$  
- **Value weights**: $\mathbf{W}^v_i \in \mathbb{R}^{d \times d_v/h}$

Where $h$ is the total number of attention heads and dimensions are split evenly across heads.

### Output Combination
Multi-head attention concatenates outputs from all heads and applies final linear transformation:
$$\text{MultiHead}(\mathbf{Q}, \mathbf{K}, \mathbf{V}) = \text{Concat}(\text{head}_1, ..., \text{head}_h)\mathbf{W}^o$$

Where $\mathbf{W}^o \in \mathbb{R}^{d_v \times d}$ is the output projection matrix.

## Key Benefits

### Representation Diversity
- Each head can learn different types of relationships
- Some heads focus on syntax, others on semantics
- Enables richer representation learning than single attention

### Computational Efficiency  
- Despite multiple heads, total computation similar to single large attention
- Dimension splitting maintains parameter count
- Parallel execution across heads

### Improved Performance
- Consistently outperforms single-head attention
- Essential component in all successful transformer variants
- Enables model to attend to information from different representation subspaces

## Implementation Details

- Typical configurations: 8, 12, or 16 heads
- Head dimension usually: $d_k = d_v = d/h$
- Maintains same total parameter count as single-head equivalent

## See also
- [attention-mechanisms](/concepts/attention-mechanisms)
- [transformer-architecture](/concepts/transformer-architecture)
- lilian-weng