---
title: Encoder-Decoder Models
category: concepts
created: 2026-12-21
updated: 2026-12-21
tags: [encoder-decoder, sequence-to-sequence, nmt, transformer-architecture, vanilla-transformer, translation, conditional-generation]
sources: [raw/feeds/2026-06-11-the-transformer-family-version-2-0.md]
confidence: high
---
# Encoder-Decoder Models
Architectural pattern commonly used in sequence-to-sequence tasks where the model must transform one sequence into another potentially different sequence. The canonical design of the [vanilla-transformer](/concepts/vanilla-transformer) and foundation for many Neural Machine Translation (NMT) systems.
## Architecture Components
### Encoder Stack
**Purpose**: Process and understand the input sequence
- Multiple layers of [multi-head-attention](/concepts/multi-head-attention) and feed-forward networks
- Bidirectional attention allows each position to attend to all input positions
- Generates rich contextual representations of the entire input sequence
- Output: Sequence of encoded representations used by decoder
### Decoder Stack
**Purpose**: Generate output sequence conditioned on encoder representations
- Masked self-attention prevents future token leakage during training
- Cross-attention layers connect to encoder outputs for conditioning
- Autoregressive generation during inference (one token at a time)
- Each position can only attend to earlier positions in the output sequence
### Cross-Attention Mechanism
**Key Innovation**: Allows decoder to selectively focus on relevant input positions
- Decoder queries attend to encoder keys and values
- Enables dynamic alignment between input and output sequences
- Critical for handling different sequence lengths and structures
## Mathematical Foundation
Using [mathematical-notation-transformers](/concepts/mathematical-notation-transformers) conventions:
### Encoder Processing
For input sequence X ∈ ℝ^(L×d):
- Self-attention: A_enc = softmax(Q_enc K_enc^T / √d_k)
- Output: H_enc = A_enc V_enc ∈ ℝ^(L×d)
### Decoder Processing
For target sequence Y ∈ ℝ^(M×d):
- Masked self-attention: A_dec = softmax(Q_dec K_dec^T / √d_k) ⊙ Mask
- Cross-attention: A_cross = softmax(Q_dec K_enc^T / √d_k)
- Combined output incorporates both decoder context and encoder information
## Application Domains
### Neural Machine Translation
**Primary Use Case**: Translating text between languages
- Encoder processes source language sequence
- Decoder generates target language sequence
- Cross-attention learns alignment between source/target words
### Conditional Text Generation
- Summarization: Long document → concise summary
- Question Answering: Context + Question → Answer
- Code Generation: Natural language → Programming code
### Multimodal Tasks
- Image Captioning: Visual features → Text description
- Speech-to-Text: Audio features → Text transcription
- Text-to-Speech: Text → Audio features
## Advantages
### Flexibility
- Handles variable-length input and output sequences
- No constraint that input and output lengths must match
- Can learn complex alignments between sequence positions
### Conditioning
- Decoder has access to full input context through cross-attention
- Enables coherent generation based on input content
- Supports both local and global context utilization
### Training Efficiency
- Parallel processing during training (teacher forcing)
- Encoder processes entire input sequence simultaneously
- Decoder can compute all positions in parallel during training
## Limitations
### Computational Complexity
- O(L²) attention complexity for both encoder and decoder
- Cross-attention adds O(L×M) computation
- Memory requirements scale quadratically with sequence length
### Inference Speed
- Autoregressive generation requires sequential token production
- Cannot parallelize decoder inference
- Longer sequences require more inference steps
### Context Length
- Attention mechanism limits maximum sequence length
- Memory constraints become prohibitive for very long sequences
- May lose information for distant dependencies
## Modern Variants
### Simplified Architectures
Following the [transformer-family-evolution](/concepts/transformer-family-evolution):
- **Encoder-only**: BERT for understanding tasks
- **Decoder-only**: [GPT](/concepts/gpt-f-diaspora) for generation tasks
- Trade architectural complexity for task-specific optimization
### Efficiency Improvements
- Sparse attention patterns to reduce complexity
- Local attention windows for long sequences
- Hierarchical attention for multi-scale processing
## Implementation Considerations
### Training Strategies
- Teacher forcing: Use ground truth tokens during training
- Scheduled sampling: Gradually introduce model predictions
- Curriculum learning: Start with shorter sequences
### Attention Patterns
- Encoder: Full bidirectional attention
- Decoder: Causal masking for self-attention
- Cross-attention: Full access to encoder outputs
### Optimization Challenges
- Gradient flow through long sequences
- Attention weight normalization
- Position encoding for different sequence lengths
## See also
- [vanilla-transformer](/concepts/vanilla-transformer)
- [transformer-architecture](/concepts/transformer-architecture)
- [multi-head-attention](/concepts/multi-head-attention)
- Neural Machine Translation
- Sequence to Sequence Learning
- Cross-Attention