ProofGen Benchmark
Benchmark suite for evaluating AI systems' ability to generate code along with formal correctness proofs. Part of the Verina codegen benchmark collection, representing a challenging test of both programming and mathematical reasoning capabilities.
Benchmark Design
Dual Capability Test: Requires systems to both generate functional code AND provide formal proofs of correctness, testing the intersection of programming and mathematical reasoning.
Verification Requirement: Goes beyond code generation to require formal verification of the generated solutions, aligning with verified-ai principles.
Challenge Level: Represents significantly higher difficulty than pure code generation benchmarks by requiring proof generation.
Performance Results
axiom-math Performance: Claims 99% success rate (187/189 problems), demonstrating exceptional capability in verified code generation.
OpenAI o3 Comparison: Last known OpenAI run achieved only 4.9% on this benchmark, highlighting the significant gap between traditional approaches and specialized verified-ai systems.
Performance Gap: The dramatic difference (99% vs 4.9%) suggests fundamental advantages of verification-focused training approaches over general language modeling.
Technical Significance
Verification Integration: Tests whether AI systems can generate both solution and proof simultaneously, rather than treating verification as post-hoc validation.
Training Signal Quality: Success on ProofGen indicates ability to generate the precise formal specifications needed for Reinforcement Learning with Verification.
Practical Applications: Performance on this benchmark directly relates to real-world applications requiring provably correct code generation.
Implications
verified-ai Validation: Strong performance supports the thesis that formal verification approaches can achieve superior results on tasks requiring mathematical precision.
Frontier Lab Gaps: Poor performance by frontier labs suggests they may not be training directly for formal verification capabilities.
Specialization Value: Indicates potential advantages of specialized approaches over general-purpose models for formal reasoning tasks.