~/wiki

Humanity's Last Exam

Confiance : high
humanity-last-examai-benchmarkknowledge-evaluationclaude-fablefallback-routingcomprehensive-assessmentbenchmark-leadership

Comprehensive benchmark designed to evaluate AI models across broad knowledge domains and reasoning capabilities. claude-fable 5 achieved 53% performance, more than 7 points ahead of the next-best model.

Performance Results

Strong performance with clear competitive advantage:

  • claude-fable 5: 53%
  • Next-best model: <46%
  • Performance gap: 7+ points

Fallback Routing Behavior

Humanity's Last Exam provides insight into claude-fable's fallback-routing system:

  • Fallback frequency: 9% of HLE tasks triggered fallback to claude-opus-48
  • Safety triggers: Likely related to cyber/bio/chemistry content
  • Transparent routing: Users notified when fallback occurs

Benchmark Characteristics

The name "Humanity's Last Exam" suggests:

  • Comprehensive evaluation across human knowledge domains
  • High-stakes assessment methodology
  • Potentially philosophical or existential framing
  • Broad coverage beyond narrow technical skills

Role in Intelligence Assessment

Part of broader intelligence evaluation alongside:

  • Intelligence Index: 64.9 (#1 performance)
  • GDPval-AA Elo: 1932 (agentic knowledge work)
  • AA-Omniscience: Knowledge benchmark improvements

Safety Architecture Insights

The 9% fallback rate on HLE tasks provides data on safety system activation patterns, showing that sensitive content detection occurs even in general knowledge evaluation contexts.

See also