~/wiki

Surge Platform

Mis à jour le 2025-12-28Confiance : medium
surge-platformhuman-evaluationblind-ratingmai-thinking-1claude-sonnet-46model-comparisonevaluation-platform

Human evaluation platform used for blind comparative testing of AI models. Notably used by microsoft to demonstrate mai-thinking-1's superiority over Claude-Sonnet-46 through blind human rater preferences.

Evaluation Methodology

Blind Rating: Human evaluators assess model outputs without knowing which model generated them Comparative Analysis: Head-to-head model performance assessment Quality Metrics: Overall preference scoring across diverse tasks

Microsoft Usage

Used to validate mai-thinking-1 performance, with blind human raters preferring it overall to Claude-Sonnet-46 - providing independent validation of the model's capabilities beyond automated benchmarks.

See also