~/wiki

GDPval-AA

Confiance : high
gdpval-aaartificial-analysisagentic-evaluationreal-world-knowledge-workelo-ratingclaude-fableperformance-measurement

Specialized evaluation metric developed by artificial-analysis for measuring AI model performance on agentic, real-world knowledge work tasks. claude-fable 5 achieved an Elo rating of 1932, ranking #1 on this benchmark.

Evaluation Focus

GDPval-AA specifically targets:

  • Agentic capabilities: Multi-step reasoning and planning
  • Real-world knowledge work: Practical business and research tasks
  • Complex problem-solving: Beyond simple question-answering
  • Long-horizon task execution: Extended reasoning chains

Claude Fable 5 Performance

  • Elo Rating: 1932
  • Ranking: #1 position
  • Task Type: Agentic real-world knowledge work

Significance in AI Evaluation

GDPval-AA represents the shift toward evaluating AI models on:

  • Practical workplace applications
  • Multi-step task completion
  • Real-world scenario handling
  • Agentic behavior assessment

This aligns with claude-fable 5's positioning as a model optimized for long-horizon-ai-tasks and complex workflows.

See also