~/wiki

LLM Performance

Mis à jour le 2026-06-11Confiance : medium
performance-metricstokens-per-secondinference-speedbenchmarkingoptimization

Metrics and techniques for measuring and optimizing the operational performance of large language models, particularly focusing on inference speed and throughput.

Key Metrics

Tokens per Second

  • Primary measure of generation speed
  • DiffusionGemma: 500+ tokens/second
  • Previous Google Gemini Diffusion: 857 tokens/second
  • Critical for real-time applications

Response Latency

  • Time to first token
  • Total generation time for complete responses
  • Important for user experience

Performance Factors

Architecture

  • Diffusion-based models showing competitive speeds
  • Model size vs. speed trade-offs
  • Hardware optimization considerations

Infrastructure

  • Cloud API hosting (NVIDIA NIM)
  • GPU acceleration requirements
  • Memory and compute resources

Benchmarking Context

Real-world performance testing provides practical insights beyond theoretical capabilities, as demonstrated by Simon Willison's hands-on testing of DiffusionGemma generation speeds.

See also