~/wiki

SWE-Bench Pro

Confiance : high
swe-bench-prosoftware-engineering-benchmarkcoding-evaluationai-coding-assessmentclaude-fablebenchmark-leadershipprogramming-tasks

Advanced software engineering benchmark used to evaluate AI models' coding capabilities on real-world programming tasks. claude-fable 5 achieved 80.3% performance compared to GPT-5.5's 58.6%, representing a significant 21.7 point advantage.

Benchmark Characteristics

SWE-Bench Pro appears to test comprehensive software engineering capabilities including:

  • Complex debugging and problem-solving
  • Multi-file code understanding and modification
  • Real-world software engineering workflows
  • Integration with existing codebases

Performance Significance

The large performance gap between claude-fable 5 (80.3%) and the next-best model (58.6%) suggests significant architectural or training improvements specifically for software engineering tasks. This aligns with anthropic's emphasis on coding capabilities in their mythos-class-models.

Role in Benchmark Leadership

SWE-Bench Pro results contribute to claude-fable's comprehensive benchmark-leadership across coding-focused evaluations, alongside cursor-bench, frontiercode-diamond, and terminal-bench.

See also