~/wiki

AA-AgentPerf

Mis à jour le 2025-12-30Confiance : high
aa-agentperfartificial-analysisagentic-inferencebenchmarkagents-per-megawattkv-cache-reusespeculative-decodingprefill-decode-disaggregationproduction-optimizationdeepseek-v4-progb300b300hopperamdpower-efficiencylong-horizon-trajectoriesinfrastructure-evaluation

Advanced benchmark developed by artificial-analysis specifically designed to evaluate agentic inference performance using long-horizon coding trajectories with production-level optimizations. Represents a significant shift from traditional throughput metrics to power-normalized deployable agent capabilities.

Key Innovation

Agents per Megawatt

Primary metric that measures power-normalized agent throughput rather than raw tokens per second:

  • Focus: Real-world deployment efficiency
  • Scope: Long-horizon coding tasks requiring multiple interaction rounds
  • Optimization: Production-ready inference techniques

Technical Features

Production Optimizations

  • KV cache reuse: Efficient memory management for extended conversations
  • Speculative decoding: Faster generation through prediction lookahead
  • Prefill/decode disaggregation: Separated processing stages for optimal resource utilization

Evaluation Methodology

  • Long-horizon trajectories: Multi-step coding tasks requiring sustained agent behavior
  • Real-world scenarios: Tasks representative of actual deployment use cases
  • Hardware-aware metrics: Power consumption integrated into performance measurement

Early Results (June 2026)

Hardware Performance

DeepSeek V4 Pro testing showed:

  • GB300 and B300: Superior agents-per-megawatt performance
  • Hopper architecture: Lower efficiency in tested configurations
  • AMD systems: Competitive but trailing NVIDIA solutions

Industry Significance

Paradigm Shift

AA-AgentPerf represents evolution from academic benchmarking toward practical deployment metrics:

  • Beyond TPS: Power-normalized rather than raw throughput focus
  • Agent-centric: Long-horizon behavior rather than single-turn generation
  • Production-ready: Incorporates real deployment optimizations

Infrastructure Impact

The benchmark addresses critical concerns for production AI deployment:

  • Cost optimization: Power efficiency directly impacts operational expenses
  • Scalability: Agent-per-watt metrics inform capacity planning
  • Hardware selection: Guides infrastructure investment decisions

See also