DeepSWE
Software engineering benchmark developed by datacurve designed to replace SWE-Bench Pro in coding evaluations by writing tasks from scratch rather than using existing repository issues. Adopted by artificial-analysis in their Coding Agent Index to reduce benchmark-gaming through repository history leakage.
Key Innovation
Clean Task Generation: Unlike SWE-Bench Pro, which used existing GitHub issues and could be gamed through repository history analysis, deepswe generates fresh coding tasks without historical contamination. This prevents models from gaining unfair advantages through training data leakage or repository pattern recognition.
Anti-Gaming Design: The benchmark explicitly addresses the problem where models could achieve high scores on SWE-Bench Pro by exploiting repository history patterns rather than demonstrating genuine coding capability.
Impact on Rankings
When artificial-analysis replaced SWE-Bench Pro with deepswe in June 2026, it materially reshuffled the Coding Agent Index rankings:
- Claude Code + Fable 5 [max]: Entered at top with score of 77
- Codex + GPT-5.5 [xhigh]: Rose to 76, overtaking previous leaders
- Claude Code + Opus 4.8 [max]: Fell to 73 from previous top position
Evaluation Philosophy
System vs Model Evaluation: deepswe's adoption highlighted the distinction between pure model capability and complete system performance, including the quality of agent harnesses and tooling integration. The benchmark measures end-to-end coding agent performance rather than isolated model reasoning.
Benchmark Saturation Concerns: Despite being newer and harder than SWE-Bench Pro, deepswe still faces the broader challenge of benchmark saturation as models rapidly improve and evaluation tasks become increasingly gameable over time.
Industry Reception
Validation Challenges: The benchmark change sparked discussions about the fairness of API evaluations when closed providers can route, fallback, or ensemble behind the scenes, making it difficult to isolate pure model performance from system engineering quality.
Harness Quality Variable: Analysis showed significant performance differences between agent harnesses using the same underlying models, suggesting that product UX capabilities vary substantially between API vendors and open-source implementations.
See also
- benchmark-gaming
- Coding Agent Index
- artificial-analysis
- SWE-Bench
- repository-contamination