GDPval-AA
Confiance : high
gdpval-aaartificial-analysisagentic-evaluationreal-world-knowledge-workelo-ratingclaude-fableperformance-measurement
Specialized evaluation metric developed by artificial-analysis for measuring AI model performance on agentic, real-world knowledge work tasks. claude-fable 5 achieved an Elo rating of 1932, ranking #1 on this benchmark.
Evaluation Focus
GDPval-AA specifically targets:
- Agentic capabilities: Multi-step reasoning and planning
- Real-world knowledge work: Practical business and research tasks
- Complex problem-solving: Beyond simple question-answering
- Long-horizon task execution: Extended reasoning chains
Claude Fable 5 Performance
- Elo Rating: 1932
- Ranking: #1 position
- Task Type: Agentic real-world knowledge work
Significance in AI Evaluation
GDPval-AA represents the shift toward evaluating AI models on:
- Practical workplace applications
- Multi-step task completion
- Real-world scenario handling
- Agentic behavior assessment
This aligns with claude-fable 5's positioning as a model optimized for long-horizon-ai-tasks and complex workflows.
See also
- artificial-analysis
- claude-fable
- long-horizon-ai-tasks
- agentic-evaluation