LLM Evaluation Methods
Comprehensive approaches to testing and measuring language model performance, based on Hugging Face's experience evaluating 15,000 models over 3 years.
Core Evaluation Approaches
Log-likelihood Evaluations: Measure conditional probability of specific answers given prompts. Calculate probability by:
- Concatenating each choice with prompt
- Extracting logits for choice tokens only
- Applying log softmax for log-probabilities
- Summing token probabilities with length normalization
Generative Evaluations: Analyze what text models actually generate through greedy decoding or sampling strategies.
Task Formulations
Multiple Choice Format (MCF): Compare likelihood of choice indices with explicit A/B/C/D presentation in prompt (as in MMLU).
Cloze Formulation (CF): Compare choice likelihoods without providing options in prompt.
Freeform Generation (FG): Evaluate accuracy of greedy generation for given prompts. Usually too difficult for pre-training ablations but primary method for post-trained models.
Evaluation Perspectives
Model Builder Needs:
- Fast, high-signal benchmarks for ablations
- Coverage of target domains/capabilities
- Intermediate checkpoint monitoring
- Scaling law predictions from smaller models
Model User Needs:
- Benchmarks matching specific use cases
- Custom evaluation design when needed
- Top contender testing for practical selection
Critical Considerations
Model Calibration: Well-calibrated models have highest probabilities for correct answers.
Tokenization Impact: Chat templates, multilingual fairness, and model-specific token handling affect results.
Evaluation Limitations: Cannot claim definitive superiority - only performance on specific proxy tasks.
Modern Challenges
Model Evolution: Pre-2022 raw text → 2022-2025 chat models → 2025+ reasoning models require different evaluation approaches.
Reasoning Models: Need to extract outputs from thinking traces (remove content between <think> tags).
Best Practices
- Respect model-expected formats and system prompts
- Account for language-dependent token generation limits
- Consider ablation studies for design choice validation
- Maintain evaluation scope awareness and appropriate caveats