Repository Contamination
A specific form of benchmark-gaming where AI models gain unfair advantages on coding benchmarks by exploiting patterns in repository history rather than demonstrating genuine programming capability. This contamination occurs when models have been trained on or can access the same repositories used to generate benchmark tasks.
Mechanism
Historical Pattern Exploitation: Models can achieve artificially high scores by recognizing repository-specific patterns, coding styles, or issue resolution approaches from their training data, rather than solving problems through general programming reasoning.
SWE-Bench Pro Vulnerability: The original SWE-Bench Pro was particularly susceptible because it used real GitHub issues and pull requests, allowing models trained on large code corpora to potentially recognize specific repositories, maintainers, or historical fix patterns.
Detection and Mitigation
Clean Task Generation: Benchmarks like deepswe address repository contamination by writing tasks from scratch rather than using existing repository issues, eliminating the possibility of historical pattern recognition.
Repository History Leakage: Models can exploit knowledge of how specific projects typically implement features, handle bugs, or structure code, making evaluation scores less representative of general coding ability.
Impact on Evaluation
Ranking Distortion: Repository contamination can significantly inflate certain models' performance, leading to misleading leaderboard positions that don't reflect true coding capability differences.
Benchmark Evolution: The recognition of repository contamination has driven the development of cleaner evaluation methodologies that generate novel tasks without historical precedent.
Broader Implications
Training Data Transparency: Repository contamination highlights the importance of understanding what repositories and code were included in model training data to properly interpret benchmark results.
Evaluation Methodology: The phenomenon demonstrates why coding evaluations must evolve beyond using existing codebases toward generating truly novel programming challenges.
See also
- benchmark-gaming
- deepswe
- data-contamination
- SWE-Bench
- Coding Agent Index