Benchmark Gaming
Mis à jour le 2025-12-30Confiance : high
benchmark-gamingevaluationswe-benchdeepswefrontiermathdata-contaminationoverfittingleaderboard-manipulationrepository-historystatic-datasetsartificial-analysis
The practice of artificially inflating performance scores on AI evaluation benchmarks through various forms of optimization that don't reflect genuine capability improvements. This undermines the reliability of benchmarks as measures of model quality and progress.
Common Techniques
Data Contamination
- Training on benchmark test sets
- Incorporating benchmark-specific patterns during training
- Repository history leakage in coding benchmarks
Overfitting Approaches
- Excessive hyperparameter tuning for specific benchmarks
- Task-specific optimizations that don't generalize
- Reward hacking in evaluation environments
Recent Developments
SWE-Bench Pro Replacement (June 2026)
artificial-analysis replaced SWE-Bench Pro with deepswe in their Coding Agent Index due to gaming concerns:
- Problem: Repository history leakage allowed unfair advantages
- Solution: deepswe generates tasks from scratch rather than using existing repository history
- Impact: Material reshuffling of leaderboard rankings
FrontierMath Issues
frontiermath v2 revealed significant benchmark construction problems:
- 42% error rate in original problems required extensive corrections
- Score inflation after error fixes while preserving rankings
- Demonstrated fragility of static evaluation datasets
Industry Impact
Harness Quality Recognition
- Benchmark quality becoming "first-class variable" in evaluation
- Distinction between model capability and product harness capability
- Questions about fairness when closed providers can route/ensemble behind scenes
Countermeasures
- Dynamic generation: Creating fresh tasks rather than static datasets
- Multi-benchmark validation: Requiring consistent performance across multiple evaluations
- Process transparency: Open evaluation methodologies and data sources
- Regular auditing: Systematic review of benchmark construction and scoring
Detection Methods
- Cross-validation on unseen tasks
- Performance analysis across related but distinct benchmarks
- Audit trails for training data and methodology
- Community peer review of evaluation claims
The gaming problem highlights the ongoing tension between standardized evaluation and the rapid evolution of AI capabilities, requiring constant vigilance and benchmark evolution.
See also
- deepswe
- frontiermath
- Evaluation Methods
- data-contamination