~/wiki

Agent Arena

Mis à jour le 2025-01-05Confiance : high
agent-arenaevaluationreal-worldcausal-tracingleaderboardagent-assessmentin-the-wild-evaluationtreatment-effectsconfirmed-successsteerabilitybash-recoverytool-hallucination1m-sessionsarena-launchmethodologyfive-signalsproduction-telemetry

Real-world agent evaluation leaderboard based on causal tracing from over 1 million actual user sessions. Represents a paradigmatic shift from synthetic benchmarks to in-the-wild performance measurement, using treatment effect estimation rather than human preference voting.

Methodology

Causal Tracing: Uses statistical methods to estimate treatment effects of different orchestrators and harnesses across real deployment scenarios.

Five Signal Framework: Evaluation based on objective metrics rather than subjective preference:

  1. Confirmed Success: Objective task completion verification
  2. Praise vs Complaint: User satisfaction indicators from actual usage
  3. Steerability: Agent responsiveness to user guidance and corrections
  4. Bash Recovery: Ability to recover from command-line errors and failures
  5. Tool Hallucination: Accuracy in tool use and API interactions

Scale and Scope

1M+ Sessions: Evaluation based on over one million real-world agent interactions across diverse use cases.

Production Telemetry: Leverages actual deployment data rather than controlled test environments.

Treatment Effect Analysis: Statistical methodology for comparing agent performance across different configurations and contexts.

Significance

Agent Arena represents a fundamental shift in AI evaluation methodology:

  • From synthetic benchmarks to real-world performance measurement
  • From human preference voting to objective success metrics
  • From laboratory conditions to production deployment assessment
  • From static evaluations to continuous performance monitoring

Evaluation Categories

The arena evaluates agents across tool use scenarios including:

  • Web search and information retrieval
  • Filesystem operations and management
  • Bash command execution and debugging
  • Image generation and visual tasks
  • Complex multi-step task orchestration

Methodological Innovation

Beyond Preference Voting: Moves away from subjective human ratings toward objective performance metrics derived from actual usage patterns.

Real-World Validity: Addresses the gap between benchmark performance and deployed system effectiveness.

Continuous Assessment: Enables ongoing evaluation of agent improvements and regressions in production environments.

See also