~/wiki

Chat Templates

Mis à jour le 2026-04-14Confiance : high
chat-templatesmodel-formattinginstruction-tuningsystem-promptstokenization

Structured formatting systems used by instruction-tuned and chat models to organize conversations with roles, system prompts, and special tokens. Critical for proper model evaluation and performance.

Evolution of Model Formatting

Pre-2022: Raw text input/output with simple completion-based prompting.

2022-2025: Chat models with role-based formatting (system, user, assistant) using JSON-like structures.

2025+: Reasoning models adding thinking tags (<thinking>, <output>) for controlled reasoning traces.

Critical Evaluation Requirements

Template Compliance: Models perform poorly without proper chat template formatting. Must respect expected format for reliable evaluation.

System Prompts: Many models require system prompts at conversation start for optimal performance.

Tokenization Interaction: Different tokenizers behave differently with spacing and special tokens in templates. Never assume identical behavior between models.

Reasoning Model Considerations

Output Extraction: Reasoning models generate thinking traces that must be removed before processing answers (typically regex removal of content between thinking tags).

Tag Recognition: Models trained with specific reasoning tags require proper handling during inference.

Model-Specific Variations

Template Differences: Each model family may use different template structures, token patterns, and role definitions.

Special Token Sensitivity: Some models (e.g., Gemma) extremely sensitive to start-of-sentence token inclusion.

Best Practices

  1. Always verify and use model-specific chat templates
  2. Include required system prompts
  3. Handle reasoning traces appropriately
  4. Test tokenization behavior with templates before evaluation
  5. Account for model-specific special token requirements

Impact on Evaluation

Improper chat template usage can drastically reduce model performance, making it appear worse than it actually is. Essential for fair comparison between models.

See also