Tokenization
Process of converting text into tokens (small units that can be characters, sub-words, or words) for language model processing. Each token is mapped to a number, creating a vocabulary that the model can understand.
Core Methods
Byte Pair Encoding (BPE): Current standard for modern LLMs. Statistical method that creates sub-word tokens based on frequency in reference text, preserving semantic connections (e.g., "similar", "dissimilar", "similarity").
Evolution: Character-based → Word-based → Sub-word (BPE) to balance vocabulary size with semantic preservation.
Critical Evaluation Considerations
Multilingual Unfairness: BPE tokenizers trained on unbalanced data favor frequent languages (typically English). Less frequent languages get split at character level, requiring orders of magnitude more tokens for equivalent text length.
Number Tokenization: Varies significantly between models - some index 0-9 digits, others store numbers up to billions individually. Affects mathematical reasoning evaluation.
Chat Templates: Different tokenizers behave differently with spacing and special tokens. Critical for post-2022 instruction-tuned models.
Model-Specific Issues
Start/End Tokens: Some models (e.g., Gemma) extremely sensitive to start-of-sentence token inclusion.
Code Model Quirks: Often train with \n\t as single token, affecting generation stopping on \n end-of-sentence markers.
Practical Impact
Token generation limits should be language-dependent for fair multilingual evaluation. Always verify tokenizer behavior with model-specific templates before evaluation.