~/wiki

Tokenization

Mis à jour le 2026-04-14Confiance : high
tokenizationnlptext-processingbpemultilingualbyte-pair-encodingmodel-evaluation

Process of converting text into tokens (small units that can be characters, sub-words, or words) for language model processing. Each token is mapped to a number, creating a vocabulary that the model can understand.

Core Methods

Byte Pair Encoding (BPE): Current standard for modern LLMs. Statistical method that creates sub-word tokens based on frequency in reference text, preserving semantic connections (e.g., "similar", "dissimilar", "similarity").

Evolution: Character-based → Word-based → Sub-word (BPE) to balance vocabulary size with semantic preservation.

Critical Evaluation Considerations

Multilingual Unfairness: BPE tokenizers trained on unbalanced data favor frequent languages (typically English). Less frequent languages get split at character level, requiring orders of magnitude more tokens for equivalent text length.

Number Tokenization: Varies significantly between models - some index 0-9 digits, others store numbers up to billions individually. Affects mathematical reasoning evaluation.

Chat Templates: Different tokenizers behave differently with spacing and special tokens. Critical for post-2022 instruction-tuned models.

Model-Specific Issues

Start/End Tokens: Some models (e.g., Gemma) extremely sensitive to start-of-sentence token inclusion.

Code Model Quirks: Often train with \n\t as single token, affecting generation stopping on \n end-of-sentence markers.

Practical Impact

Token generation limits should be language-dependent for fair multilingual evaluation. Always verify tokenizer behavior with model-specific templates before evaluation.

See also