~/wiki

Benchmark Gaming

Mis à jour le 2025-12-30Confiance : high
benchmark-gamingevaluationswe-benchdeepswefrontiermathdata-contaminationoverfittingleaderboard-manipulationrepository-historystatic-datasetsartificial-analysis

The practice of artificially inflating performance scores on AI evaluation benchmarks through various forms of optimization that don't reflect genuine capability improvements. This undermines the reliability of benchmarks as measures of model quality and progress.

Common Techniques

Data Contamination

  • Training on benchmark test sets
  • Incorporating benchmark-specific patterns during training
  • Repository history leakage in coding benchmarks

Overfitting Approaches

  • Excessive hyperparameter tuning for specific benchmarks
  • Task-specific optimizations that don't generalize
  • Reward hacking in evaluation environments

Recent Developments

SWE-Bench Pro Replacement (June 2026)

artificial-analysis replaced SWE-Bench Pro with deepswe in their Coding Agent Index due to gaming concerns:

  • Problem: Repository history leakage allowed unfair advantages
  • Solution: deepswe generates tasks from scratch rather than using existing repository history
  • Impact: Material reshuffling of leaderboard rankings

FrontierMath Issues

frontiermath v2 revealed significant benchmark construction problems:

  • 42% error rate in original problems required extensive corrections
  • Score inflation after error fixes while preserving rankings
  • Demonstrated fragility of static evaluation datasets

Industry Impact

Harness Quality Recognition

  • Benchmark quality becoming "first-class variable" in evaluation
  • Distinction between model capability and product harness capability
  • Questions about fairness when closed providers can route/ensemble behind scenes

Countermeasures

  1. Dynamic generation: Creating fresh tasks rather than static datasets
  2. Multi-benchmark validation: Requiring consistent performance across multiple evaluations
  3. Process transparency: Open evaluation methodologies and data sources
  4. Regular auditing: Systematic review of benchmark construction and scoring

Detection Methods

  • Cross-validation on unseen tasks
  • Performance analysis across related but distinct benchmarks
  • Audit trails for training data and methodology
  • Community peer review of evaluation claims

The gaming problem highlights the ongoing tension between standardized evaluation and the rapid evolution of AI capabilities, requiring constant vigilance and benchmark evolution.

See also