---
title: Transformer Architecture
category: concepts
created: 2026-06-11
updated: 2026-12-20
tags: [transformer-architecture, attention-mechanisms, neural-networks, llm-foundation, self-attention, encoder-decoder, positional-encoding, inference-optimization, vanilla-transformer, mathematical-notation]
sources: [raw/feeds/2026-06-11-transformer-architecture.md, raw/feeds/2026-06-11-large-transformer-model-inference-optimization.md, raw/feeds/2026-06-11-the-transformer-family-version-2-0.md]
confidence: high
---
# Transformer Architecture
The foundational neural network architecture that revolutionized natural language processing and became the backbone of modern large language models. Introduced the self-attention mechanism and parallel processing capabilities that enabled scaling to massive model sizes.
## Core Components
### Vanilla Transformer (Vaswani et al., 2017)
The original "vanilla Transformer" established the encoder-decoder architecture commonly used in Neural Machine Translation (NMT) models. This foundational design was later simplified into:
- **Encoder-only**: BERT and similar models for understanding tasks
- **Decoder-only**: GPT and similar models for generation tasks
### Mathematical Framework
According to lilian-weng's comprehensive analysis, transformer architectures are characterized by:
- **Model dimension** ($d$): Hidden state and positional encoding size
- **Multi-head attention** ($h$ heads): Parallel attention mechanisms
- **Sequence length** ($L$): Input sequence processing capacity
- **Layer depth** ($N$): Total attention layers (excluding MoE)
### Attention Mechanism
The core innovation enabling parallel processing through [attention-mechanisms](/concepts/attention-mechanisms):
- **Query-Key-Value** matrices: $\mathbf{Q}$, $\mathbf{K}$, $\mathbf{V}$ transformations
- **Attention scores**: $\mathbf{A} = \text{softmax}(\mathbf{Q}\mathbf{K}^\top / \sqrt{d_k})$
- **Multi-head parallel processing**: Multiple attention functions running simultaneously
## Evolution and Variants
The transformer family has expanded significantly since 2017, with lilian-weng's 2023 survey documenting three years of architectural improvements that roughly doubled the scope of transformer variants.
## Key Advantages
- **Parallelization**: Unlike RNNs, all positions processed simultaneously
- **Long-range dependencies**: Attention mechanism captures distant relationships
- **Scalability**: Architecture enables training models with billions/trillions of parameters
- **Versatility**: Successful across diverse tasks (translation, generation, understanding)
## See also
- [attention-mechanisms](/concepts/attention-mechanisms)
- [multi-head-attention](/concepts/multi-head-attention)
- [positional-encoding](/concepts/positional-encoding)
- lilian-weng