Humanity's Last Exam
Confiance : high
humanity-last-examai-benchmarkknowledge-evaluationclaude-fablefallback-routingcomprehensive-assessmentbenchmark-leadership
Comprehensive benchmark designed to evaluate AI models across broad knowledge domains and reasoning capabilities. claude-fable 5 achieved 53% performance, more than 7 points ahead of the next-best model.
Performance Results
Strong performance with clear competitive advantage:
- claude-fable 5: 53%
- Next-best model: <46%
- Performance gap: 7+ points
Fallback Routing Behavior
Humanity's Last Exam provides insight into claude-fable's fallback-routing system:
- Fallback frequency: 9% of HLE tasks triggered fallback to claude-opus-48
- Safety triggers: Likely related to cyber/bio/chemistry content
- Transparent routing: Users notified when fallback occurs
Benchmark Characteristics
The name "Humanity's Last Exam" suggests:
- Comprehensive evaluation across human knowledge domains
- High-stakes assessment methodology
- Potentially philosophical or existential framing
- Broad coverage beyond narrow technical skills
Role in Intelligence Assessment
Part of broader intelligence evaluation alongside:
- Intelligence Index: 64.9 (#1 performance)
- GDPval-AA Elo: 1932 (agentic knowledge work)
- AA-Omniscience: Knowledge benchmark improvements
Safety Architecture Insights
The 9% fallback rate on HLE tasks provides data on safety system activation patterns, showing that sensitive content detection occurs even in general knowledge evaluation contexts.
See also
- benchmark-leadership
- claude-fable
- fallback-routing
- intelligence-assessment
- knowledge-evaluation