~/wiki

ProofGen Benchmark

Mis à jour le 2026-01-03Confiance : high
proofgen-benchmarkverina-codegenformal-verificationcode-generationcorrectness-proofsmathematical-reasoningbenchmarkaxiom-mathperformance-comparison

Benchmark suite for evaluating AI systems' ability to generate code along with formal correctness proofs. Part of the Verina codegen benchmark collection, representing a challenging test of both programming and mathematical reasoning capabilities.

Benchmark Design

Dual Capability Test: Requires systems to both generate functional code AND provide formal proofs of correctness, testing the intersection of programming and mathematical reasoning.

Verification Requirement: Goes beyond code generation to require formal verification of the generated solutions, aligning with verified-ai principles.

Challenge Level: Represents significantly higher difficulty than pure code generation benchmarks by requiring proof generation.

Performance Results

axiom-math Performance: Claims 99% success rate (187/189 problems), demonstrating exceptional capability in verified code generation.

OpenAI o3 Comparison: Last known OpenAI run achieved only 4.9% on this benchmark, highlighting the significant gap between traditional approaches and specialized verified-ai systems.

Performance Gap: The dramatic difference (99% vs 4.9%) suggests fundamental advantages of verification-focused training approaches over general language modeling.

Technical Significance

Verification Integration: Tests whether AI systems can generate both solution and proof simultaneously, rather than treating verification as post-hoc validation.

Training Signal Quality: Success on ProofGen indicates ability to generate the precise formal specifications needed for Reinforcement Learning with Verification.

Practical Applications: Performance on this benchmark directly relates to real-world applications requiring provably correct code generation.

Implications

verified-ai Validation: Strong performance supports the thesis that formal verification approaches can achieve superior results on tasks requiring mathematical precision.

Frontier Lab Gaps: Poor performance by frontier labs suggests they may not be training directly for formal verification capabilities.

Specialization Value: Indicates potential advantages of specialized approaches over general-purpose models for formal reasoning tasks.

See also