~/wiki

FrontierCode

Mis à jour le 2025-01-05Confiance : high
frontiercodebenchmarkcode-qualityevaluationcognitionmergeable-codemaintainabilityslopswe-benchfalse-positivesreal-world-assessmentopen-source-maintainersregression-safetycleanlinessscope-correctnesswar-on-slop40-hour-tasksopus-4.8-performancemetr-validation

Advanced coding evaluation benchmark developed by cognition that measures whether generated code is actually mergeable and maintainable, not just test-passing. Represents a significant evolution beyond SWE-Bench toward real-world code quality assessment and direct response to the "war on slop" in AI-generated code.

Methodology

FrontierCode tasks are built with open-source maintainers, with each task requiring 40+ hours of work and evaluated across multiple dimensions:

  • Regression safety: Does the code break existing functionality?
  • Cleanliness: Is the code well-structured and readable?
  • Scope correctness: Does the implementation address the right problem boundaries?
  • Test correctness: Are tests meaningful and comprehensive?
  • Maintainability: Can the code be reasonably maintained over time?

Performance Results

The benchmark reveals significantly lower performance than traditional coding evaluations:

  • Best performing model (Opus 4.8) scores only ~13% on the hardest subset
  • Contrasts sharply with 50%+ scores common on SWE-Bench-style evaluations
  • Suggests coding capabilities are much less "solved" than popular benchmarks imply

Context and Motivation

FrontierCode was explicitly inspired by frontiermath and addresses identified problems with existing benchmarks:

  • metr findings that many SWE-Bench-passing PRs would not be merged into main
  • Problem of false positive trajectories in benchmark evaluation
  • Need for better articulation of code quality rubrics beyond test passage

The benchmark emerges from the broader "War on Slop" - industry pushback against low-quality AI-generated code that passes tests but fails production requirements.

See also