Surge Platform
Mis à jour le 2025-12-28Confiance : medium
surge-platformhuman-evaluationblind-ratingmai-thinking-1claude-sonnet-46model-comparisonevaluation-platform
Human evaluation platform used for blind comparative testing of AI models. Notably used by microsoft to demonstrate mai-thinking-1's superiority over Claude-Sonnet-46 through blind human rater preferences.
Evaluation Methodology
Blind Rating: Human evaluators assess model outputs without knowing which model generated them Comparative Analysis: Head-to-head model performance assessment Quality Metrics: Overall preference scoring across diverse tasks
Microsoft Usage
Used to validate mai-thinking-1 performance, with blind human raters preferring it overall to Claude-Sonnet-46 - providing independent validation of the model's capabilities beyond automated benchmarks.
See also
- mai-thinking-1
- human-evaluation
- model-benchmarking