---
title: Datacurve
category: concepts
created: 2025-01-04
updated: 2025-01-04
tags: [datacurve, deepswe, benchmark-development, swe-bench-replacement, benchmark-gaming-resistance, artificial-analysis, coding-evaluation]
sources: [raw/feeds/2026-06-13--ainews-fable-and-mythos-officially-too-dangerous-to-release.md]
confidence: medium
---
# Datacurve
Organization that developed [deepswe](/concepts/deepswe), a software engineering benchmark designed to replace SWE-Bench Pro in coding evaluations. Their approach focuses on writing tasks from scratch rather than using existing repository issues to reduce [benchmark-gaming](/concepts/benchmark-gaming) through repository history contamination.
## Key Innovation
### DeepSWE Development
- **Task generation**: Creates coding challenges from scratch instead of mining existing repositories
- **Contamination resistance**: Prevents models from gaming through repository history access
- **Evaluation integrity**: Maintains benchmark validity as models become more sophisticated
### Industry Adoption
artificial-analysis adopted [deepswe](/concepts/deepswe) in their Coding Agent Index, replacing SWE-Bench Pro due to gaming concerns. This change materially reshuffled rankings:
- Claude Code + claude-fable 5 entered at top with score of 77
- Codex + GPT-5.5 rose to 76
- Demonstrated impact of benchmark methodology on evaluation outcomes
## Benchmark Philosophy
Represents shift toward:
- **Ground-up task creation** rather than mining existing work
- **Gaming-resistant design** that maintains validity over time
- **Realistic evaluation** of coding capabilities without dataset leakage
- **Methodological transparency** in benchmark construction
## See also
- [deepswe](/concepts/deepswe)
- [benchmark-gaming](/concepts/benchmark-gaming)
- artificial-analysis
- Coding Agent Index
- [SWE-Bench](/concepts/swe-bench-pro)