Chat Templates
Structured formatting systems used by instruction-tuned and chat models to organize conversations with roles, system prompts, and special tokens. Critical for proper model evaluation and performance.
Evolution of Model Formatting
Pre-2022: Raw text input/output with simple completion-based prompting.
2022-2025: Chat models with role-based formatting (system, user, assistant) using JSON-like structures.
2025+: Reasoning models adding thinking tags (<thinking>, <output>) for controlled reasoning traces.
Critical Evaluation Requirements
Template Compliance: Models perform poorly without proper chat template formatting. Must respect expected format for reliable evaluation.
System Prompts: Many models require system prompts at conversation start for optimal performance.
Tokenization Interaction: Different tokenizers behave differently with spacing and special tokens in templates. Never assume identical behavior between models.
Reasoning Model Considerations
Output Extraction: Reasoning models generate thinking traces that must be removed before processing answers (typically regex removal of content between thinking tags).
Tag Recognition: Models trained with specific reasoning tags require proper handling during inference.
Model-Specific Variations
Template Differences: Each model family may use different template structures, token patterns, and role definitions.
Special Token Sensitivity: Some models (e.g., Gemma) extremely sensitive to start-of-sentence token inclusion.
Best Practices
- Always verify and use model-specific chat templates
- Include required system prompts
- Handle reasoning traces appropriately
- Test tokenization behavior with templates before evaluation
- Account for model-specific special token requirements
Impact on Evaluation
Improper chat template usage can drastically reduce model performance, making it appear worse than it actually is. Essential for fair comparison between models.