~/wiki

LLM Evaluation Methods

Mis à jour le 2026-04-14Confiance : high
llm-evaluationmodel-testingbenchmarksevaluation-methodslog-likelihoodgenerative-evaluationmodel-calibration

Comprehensive approaches to testing and measuring language model performance, based on Hugging Face's experience evaluating 15,000 models over 3 years.

Core Evaluation Approaches

Log-likelihood Evaluations: Measure conditional probability of specific answers given prompts. Calculate probability by:

  1. Concatenating each choice with prompt
  2. Extracting logits for choice tokens only
  3. Applying log softmax for log-probabilities
  4. Summing token probabilities with length normalization

Generative Evaluations: Analyze what text models actually generate through greedy decoding or sampling strategies.

Task Formulations

Multiple Choice Format (MCF): Compare likelihood of choice indices with explicit A/B/C/D presentation in prompt (as in MMLU).

Cloze Formulation (CF): Compare choice likelihoods without providing options in prompt.

Freeform Generation (FG): Evaluate accuracy of greedy generation for given prompts. Usually too difficult for pre-training ablations but primary method for post-trained models.

Evaluation Perspectives

Model Builder Needs:

  • Fast, high-signal benchmarks for ablations
  • Coverage of target domains/capabilities
  • Intermediate checkpoint monitoring
  • Scaling law predictions from smaller models

Model User Needs:

  • Benchmarks matching specific use cases
  • Custom evaluation design when needed
  • Top contender testing for practical selection

Critical Considerations

Model Calibration: Well-calibrated models have highest probabilities for correct answers.

Tokenization Impact: Chat templates, multilingual fairness, and model-specific token handling affect results.

Evaluation Limitations: Cannot claim definitive superiority - only performance on specific proxy tasks.

Modern Challenges

Model Evolution: Pre-2022 raw text → 2022-2025 chat models → 2025+ reasoning models require different evaluation approaches.

Reasoning Models: Need to extract outputs from thinking traces (remove content between <think> tags).

Best Practices

  • Respect model-expected formats and system prompts
  • Account for language-dependent token generation limits
  • Consider ablation studies for design choice validation
  • Maintain evaluation scope awareness and appropriate caveats

See also