~/wiki

Concepts — vue longue

retour à la liste

Toutes les pages concaténées sur un seul document, pour un Ctrl-F direct.

Baseten Partnership

page dédiée →

Strategic partnership between microsoft and baseten providing enterprise-controlled fine-tuning for mai-models with "100% eyes-off" data privacy guarantees.

Key Features

100% Eyes-Off: Complete data privacy - no human access to enterprise training data Enterprise-Controlled Fine-tuning: Customer maintains full control over model customization clean-data-lineage: Maintained throughout the fine-tuning process Privacy Compliance: Addresses enterprise data governance requirements

Strategic Value

Enterprise Adoption: Removes data privacy barriers for large organization AI deployment Competitive Advantage: Differentiates MAI models from competitors on privacy grounds Market Positioning: Positions Microsoft as enterprise-first AI provider

Technical Implementation

Built on mai-thinking-1 and broader MAI model family, enabling organizations to create specialized versions while maintaining data confidentiality and regulatory compliance.

See also

Chinchilla Optimal

page dédiée →

Training methodology that balances model parameters and training tokens to achieve compute-optimal performance, based on scaling law research. Referenced in mai-thinking-1 development as a baseline for ablation studies.

MAI-Thinking-1 Implementation

Microsoft conducted ablations at roughly 100-200 tokens per parameter, described as "around Chinchilla optimal" for their setup, though noting differences from dense-model heuristics due to moe-architecture structure.

Scaling Law Foundation

Based on research demonstrating optimal compute allocation between:

  • Model parameter count
  • Training token quantity
  • Computational budget constraints

MoE Considerations

Traditional Chinchilla optimal ratios may require adjustment for Mixture of Experts architectures due to different parameter utilization patterns and effective model capacity calculations.

Strategic Implications

Chinchilla-optimal training enables:

  • Maximum performance for given compute budget
  • Efficient resource allocation decisions
  • Comparative evaluation of architecture efficiency
  • Foundation for scaling law extrapolation

Research Impact

Established fundamental principles for compute-efficient training that influence model development decisions across the industry, providing scientific basis for training resource allocation.

See also

Clean Data Lineage

page dédiée →

Training methodology emphasizing transparent, traceable data sources without third-party model distillation or synthetic data generation. Pioneered by microsoft in the mai-models family, particularly mai-thinking-1.

Core Principles

No Distillation: Zero use of outputs from third-party models during training No Synthetic Data: Reliance on authentic, naturally-occurring data sources
Transparent Sources: Clear documentation of all data origins and processing steps Quality Control: Rigorous extraction, deduplication, and curation processes

Microsoft's Implementation

Data Sources:

  • common-crawl web data
  • Private, curated datasets
  • Targeted sub-pipelines for different domains

Quality Assurance:

  • Heavy extraction and deduplication work
  • DSPy-GEPA optimized LLM judges for quality scoring
  • Domain-specific curation pipelines

Enterprise Value

Trust: Clear data provenance for compliance and auditing Control: No dependency on competitor model outputs Quality: Higher signal-to-noise ratio through careful curation Legal Safety: Reduced IP and licensing complications

Industry Impact

Represents pushback against widespread use of synthetic data and model distillation, emphasizing the value of authentic data sources for frontier model development.

See also

Framework for optimizing language model pipelines through systematic prompt engineering and quality assessment. Notably used by microsoft in mai-thinking-1 development for advanced data curation and quality scoring through optimized LLM judges.

Core Concepts

DSPy provides systematic approaches to:

  • Prompt Optimization: Automatic improvement of prompts through optimization algorithms
  • Pipeline Composition: Chaining multiple LLM operations with optimization
  • Quality Assessment: Using optimized LLM judges for data scoring and curation

GEPA Integration

GEPA (Generalized Evaluation and Prompt Adaptation) represents an advanced DSPy application involving:

  • Late-interaction techniques for efficient retrieval and scoring
  • Optimized LLM judges for quality assessment
  • Systematic prompt adaptation based on performance metrics

Microsoft MAI Implementation

In mai-thinking-1 development, Microsoft leveraged DSPy for:

Data Curation

  • Quality Scoring: DSPy-optimized LLM judges evaluated training data quality
  • Domain-Specific Pipelines: Targeted curation for different content domains
  • Extraction and Deduplication: Systematic processing of Common Crawl and private sources

Training Pipeline Integration

  • Pre-training Data: Quality scoring during data preparation
  • Scaling Decisions: Optimization metrics for architecture promotion
  • Efficiency Measurement: Systematic evaluation of training effectiveness

Research Community Impact

The disclosure of DSPy usage in MAI-Thinking-1 development generated significant attention from the late-interaction and optimization research communities, demonstrating practical applications of systematic prompt optimization at frontier model scale.

Technical Applications

LLM Judge Optimization

  • Automatic improvement of evaluation prompts
  • Consistency optimization across evaluation runs
  • Domain-specific judge adaptation

Pipeline Orchestration

  • Multi-stage processing with optimized transitions
  • Error correction and quality gating
  • Performance monitoring and adaptation

See also

Efficiency Gain Metric

page dédiée →

Metric used by microsoft in mai-thinking-1 development for making architecture promotion decisions during the scaling-ladder process. Measures how much extra compute the baseline architecture would need to match a candidate architecture's loss performance.

Application in MAI-Thinking-1

Architecture Comparison Framework

The Efficiency Gain metric provides quantitative basis for comparing different architectural choices by measuring their relative computational efficiency for achieving equivalent loss performance.

Promotion Decision Criteria

Used as primary metric for deciding which candidate architectures to promote to larger scales during the systematic scaling ladder evaluation process.

Compute Resource Optimization

Enables data-driven decisions about resource allocation by quantifying the computational cost difference between architectural alternatives.

Technical Implementation

Baseline vs Candidate Evaluation

  • Baseline Architecture: Reference architecture with known compute requirements
  • Candidate Architecture: Alternative architecture being evaluated
  • Efficiency Calculation: Ratio of compute needed by baseline to match candidate's loss

Loss Performance Matching

The metric specifically measures compute requirements to achieve equivalent loss performance, rather than other metrics that might not directly correlate with training efficiency.

Scale-Aware Assessment

Applied across different compute scales during the scaling ladder process, ensuring architectural decisions remain optimal at various training scales.

Research Significance

Systematic Architecture Selection

Provides objective framework for architecture selection in large-scale model development, moving beyond intuitive or experience-based decisions.

Resource Planning

Enables predictive planning for computational resource requirements when scaling architectural decisions to full model training.

Reproducible Methodology

Creates standardized approach for architectural evaluation that can be applied systematically across different model development projects.

See also

Advanced technique used in conjunction with dspy for optimizing LLM judges in data curation and quality scoring. Notably employed by microsoft in mai-thinking-1 development for pretraining data quality assessment.

Integration with DSPy

GEPA works within the DSPy framework to enhance LLM judge optimization, particularly for:

  • Pretraining data curation
  • Quality scoring of training examples
  • Automated data pipeline evaluation
  • Late-interaction optimization

MAI-Thinking-1 Implementation

Microsoft's use of DSPy-optimized LLM judges with GEPA represented a sophisticated approach to data quality control, contributing to the model's clean data lineage and high performance outcomes.

Technical Community Interest

Generated significant attention from the DSPy and late-interaction research communities, highlighting the growing importance of optimized evaluation systems in frontier model development.

Relationship to Data Quality

Part of Microsoft's broader emphasis on clean-data-lineage, demonstrating how advanced curation techniques can substitute for synthetic data or distillation approaches while maintaining high model performance.

See also

microsoft's internally developed language model series emphasizing clean lineage, exceptional data quality, and hill climbing capabilities. Designed to enable companies to build their own specialist models rather than relying solely on generalist models. At Build 2026, Microsoft announced seven new MAI models demonstrating competitive frontier capabilities and unprecedented technical-transparency.

Design Philosophy

Clean Lineage Foundation

Starting with pre-training using very high data quality with extensive ablation studies. satya-nadella emphasized this is "becoming even harder to build a clean lineage model just because there's so much stuff out there that you truly need to ablate out to be able to have a fantastic pre-trained model."

This addresses a key limitation of many open weight models that "look great on one benchmark or two, but they're not great on practice."

Cognitive Core Pursuit

Central to MAI development is pursuing the "cognitive-core" - fundamental intelligence patterns that can serve as the foundation for specialized capabilities. This approach prioritizes essential intelligence over pure scale.

Hill Climbing Architecture

Scaffold System

MAI models include a "hill climb scaffold" enabling customers to:

  • Build specialist models from the generalist foundation
  • Implement trace-collection for continuous improvement
  • Develop private-evals specific to their domain
  • Create proprietary intellectual property through model specialization

Temporal Scaffolding Innovation

Demonstrated through the land-o-lakes-demo where:

  • GPT-55 was used to collect traces
  • A 5B reasoning model achieved higher performance using those traces
  • This represents "a new frontier" in AI capability development

Platform Integration Strategy

MAI models serve as the foundation for Microsoft's frontier-intelligence-platform approach:

  • Enable "first-class participants" who can point to AI they created
  • Support enterprise specialization rather than generic AI consumption
  • Integrate with multi-model harnesses like openclaw and scout
  • Connect with enterprise context through work-iq

Seven Model Family (Build 2026)

Microsoft announced seven new MAI models demonstrating:

  • Competitive frontier capabilities
  • Unprecedented technical transparency
  • Specialized capabilities across different domains
  • Support for enterprise-controlled fine-tuning

Training Strategy Advantages

Data Quality Focus

  • Extensive ablation studies to ensure clean training data
  • Careful curation to avoid contamination common in open models
  • Focus on quality over quantity in training corpus

Specialized Development Path

  • Not just generalist models but foundation for specialization
  • Enables customers to develop proprietary AI capabilities
  • Supports enterprise-specific use cases and requirements

Competitive Positioning

MAI models position Microsoft uniquely as:

  • Both platform provider and frontier model developer
  • Enabling customer AI development rather than just AI consumption
  • Balancing technical capability with ecosystem enablement
  • Addressing practical deployment challenges through clean architecture

See also

microsoft's custom AI chip optimized for running mai-models, delivering significant performance and efficiency improvements over standard GB200-GPUs for Microsoft's model inference workloads.

Performance Characteristics

Cost Efficiency: 30% better performance per dollar compared to GB200 Power Efficiency: 1.4x performance-per-watt gain versus GB200 Optimization Target: End-to-end MAI model serving and inference

Strategic Importance

Hardware-Software Co-design: Custom silicon optimized specifically for MAI model architectures Cost Advantage: Significant operational cost benefits for Microsoft's AI services Competitive Moat: Hardware optimization as differentiation strategy

Technical Integration

Optimized for mai-thinking-1 and broader MAI family serving, representing Microsoft's investment in full-stack AI infrastructure control from silicon to software.

See also

MFU Disclosure

page dédiée →

Model FLOPs Utilization (MFU) metrics disclosure representing the percentage of theoretical hardware performance achieved during training. microsoft's disclosure of exact MFU numbers across iterations for mai-thinking-1 was noted as unprecedented transparency for frontier model development.

Significance of Disclosure

MFU numbers are rarely shared at frontier model scale because they reveal:

  • Infrastructure efficiency and capabilities
  • Engineering quality and optimization expertise
  • Competitive training cost information
  • Hardware utilization optimization techniques

Technical Importance

MFU measurements enable:

  • Objective comparison of training infrastructure efficiency
  • Identification of optimization opportunities
  • Hardware procurement and scaling decisions
  • Engineering team performance assessment

Microsoft's Transparency

The disclosure of exact MFU across training iterations demonstrated technical-transparency that multiple researchers highlighted as "rarely shared at this scale," contributing to the research community's positive reception of the technical report.

Industry Impact

Such detailed disclosure sets new standards for frontier model transparency and provides valuable reference points for the broader AI research community working on training efficiency optimization.

See also

MoE Architecture

page dédiée →

Mixture of Experts (MoE) is a neural network architecture that uses multiple specialized sub-networks (experts) with a gating mechanism to route inputs to the most relevant experts. Enables efficient scaling by activating only a subset of total parameters for each inference, as demonstrated in mai-thinking-1 and other frontier models.

Core Concepts

Expert Specialization:

  • Multiple specialized sub-networks within single model
  • Each expert develops domain-specific capabilities
  • Gating network learns to route inputs to appropriate experts
  • Enables model specialization without full parameter activation

Parameter Efficiency:

  • Total Parameters: Full model size including all experts
  • Active Parameters: Subset activated for any given input
  • Example: mai-thinking-1 has 1T total parameters but only 35B active
  • Significant compute savings during inference while maintaining model capacity

Scaling Advantages

Compute Efficiency:

  • Linear scaling of experts with sub-linear compute growth
  • Better performance per active parameter compared to dense models
  • Different scaling laws compared to traditional dense architectures
  • More efficient than equivalent dense models at same active parameter count

Specialization Benefits:

  • Experts can develop domain-specific knowledge (code, math, language-specific)
  • Reduced interference between different capability areas
  • Better performance on diverse tasks within single model
  • Easier to add new capabilities through additional experts

Implementation Challenges

Training Complexity:

  • Load balancing across experts to prevent expert collapse
  • Routing efficiency and stability during training
  • More complex infrastructure requirements
  • Different scaling heuristics compared to dense models

Architecture Design:

  • Optimal expert count and size determination
  • Gating mechanism design and training
  • Expert specialization encouragement vs generalization
  • Memory and communication overhead management

Microsoft MAI Implementation

MAI-Thinking-1 Specifications:

  • 1T total parameters with 35B active parameters
  • ~28x parameter efficiency ratio
  • 256K context window maintained across all experts
  • Optimized for MAIA 200 custom hardware

MAI-Code-1-Flash:

  • 137B total parameters with 5B active parameters
  • ~27x parameter efficiency for coding tasks
  • Achieves 51% SWE-Bench Pro despite small active footprint
  • Specialized for VS Code and GitHub Copilot integration

Training Methodology

Scaling Decisions:

  • Efficiency Gain (EG) metrics for architecture promotion
  • Ablations around Chinchilla-optimal token ratios adapted for MoE
  • Private evaluation sets for expert specialization assessment
  • Custom scaling laws for MoE vs dense model comparison

Expert Development:

  • Domain-specific data routing during training
  • Load balancing mechanisms to ensure expert utilization
  • Specialization encouragement through routing policies
  • Quality control across expert outputs

Hardware Optimization

Custom Silicon Integration:

  • MAIA 200 optimization providing 30% better performance per dollar
  • 1.4x performance-per-watt gain for MAI models
  • Hardware-software co-design for expert routing efficiency
  • Memory hierarchy optimization for sparse activation patterns

Industry Context

Competitive Landscape:

  • Most frontier models now use some form of MoE architecture
  • Enables larger models without proportional compute increase
  • Critical for cost-effective serving of large language models
  • Standard approach for balancing capability with efficiency

Future Directions:

  • Increasing expert counts and specialization
  • Dynamic expert creation and pruning
  • Cross-modal expert architectures
  • More sophisticated routing mechanisms

MoE architecture represents a fundamental shift toward sparse, efficient scaling in language models, enabling the parameter counts necessary for frontier performance while maintaining practical deployment constraints.

See also

Private NLL Evaluation

page dédiée →

Internal evaluation methodology using Negative Log Likelihood (NLL) on private datasets for making scaling and architecture decisions during model development. Employed by microsoft in mai-thinking-1 development for systematic model progression.

Data Composition for MAI-Thinking-1

Microsoft's private NLL evaluation set comprised:

  • 50% code
  • 17.5% STEM
  • 17.5% math
  • 10% general knowledge
  • 5% multilingual

Role in Scaling Decisions

Used to evaluate candidate architectures during the scaling-ladder process, providing consistent performance measurement across different model scales and configurations.

Technical Implementation

Negative Log Likelihood provides a fundamental loss measurement that enables:

  • Objective comparison between model architectures
  • Scaling law analysis and extrapolation
  • Data-driven decisions on architecture promotion
  • Consistent evaluation across training iterations

Strategic Advantage

Private evaluation sets enable companies to make scaling decisions based on proprietary benchmarks that may better reflect target use cases than public benchmarks, while maintaining evaluation consistency across development cycles.

See also

Scaling Ladder

page dédiée →

Training methodology used by microsoft in developing mai-thinking-1, involving systematic architecture evaluation and promotion decisions based on performance metrics across different compute scales.

Core Methodology

Architecture Evaluation Process

The scaling ladder involves testing candidate architectures at smaller scales before committing to full-scale training, allowing for data-driven decisions about which architectures to promote to larger scales.

Efficiency Gain Metric

Architecture promotion decisions are based on the efficiency-gain-metric, which quantifies how much extra compute the baseline architecture would need to match a candidate architecture's loss performance.

Systematic Scaling Decisions

Rather than intuitive or heuristic-based scaling choices, the methodology provides quantitative framework for architecture selection at each scale tier.

Application in MAI-Thinking-1

Ablation Studies

Conducted ablations at approximately 100/200 tokens per parameter, described as "Chinchilla optimal" for the MoE setup, though differing from dense model heuristics.

Data-Driven Promotion

Architecture candidates systematically evaluated and promoted based on performance metrics rather than subjective assessment or industry conventions.

Scale-Aware Optimization

Methodology accounts for the fact that optimal architectures may differ at various compute scales, particularly for MoE configurations.

Technical Innovation

Beyond Dense Model Heuristics

The scaling ladder methodology explicitly accounts for MoE architectural differences, recognizing that traditional dense model scaling laws may not apply directly.

Systematic Experimentation

Provides framework for rigorous experimental methodology in large-scale model development, moving beyond ad-hoc scaling decisions.

Resource Optimization

Enables efficient use of computational resources by making informed decisions about architecture promotion rather than training all candidates to full scale.

Research Community Impact

The detailed disclosure of scaling ladder methodology in Microsoft's technical report provides actionable framework for other researchers developing large-scale models with systematic architecture evaluation.

See also

Scaling Ladder Methodology

page dédiée →

Systematic approach to model development that uses incremental scaling and architecture comparison to optimize model design before full-scale training. Prominently featured in microsoft's mai-thinking-1 development process.

Core Methodology

Efficiency Gain (EG) Metric

Architecture promotion decisions based on Efficiency Gain calculation:

  • Definition: How much extra compute the baseline would need to match the candidate's loss
  • Application: Systematic comparison of architectural variants
  • Decision Framework: Objective criteria for architecture selection

Scaling Progression

Incremental scaling through defined checkpoints:

  • Small Scale Testing: Initial architecture validation
  • Progressive Scaling: Systematic increase in model size and training data
  • Architecture Refinement: Continuous optimization based on EG metrics

Training Schedule Optimization

Ablation Studies

Conducted at approximately 100-200 tokens per parameter:

  • Chinchilla Optimal Range: Roughly optimal for the experimental setup
  • MoE Adaptations: Different from dense model heuristics due to moe-architecture
  • Resource Efficiency: Systematic testing without full-scale resource commitment

Validation Methodology

  • Internal NLL Set: Private validation dataset for scaling decisions
  • Loss Tracking: Systematic monitoring of training loss across scales
  • Performance Prediction: Extrapolation from smaller scale results

MAI-Thinking-1 Implementation

Data Composition

Internal validation set composition for scaling decisions:

  • 50% code
  • 17.5% STEM
  • 17.5% math
  • 10% general knowledge
  • 5% multilingual

Architecture Decisions

  • MoE Configuration: Optimal expert count and routing strategies
  • Parameter Allocation: Balance between active and total parameters
  • Context Window: Optimization for 256K token context length

Research Impact

The detailed disclosure of scaling ladder methodology in MAI-Thinking-1's technical report provides unprecedented insights into systematic model development, serving as a practical guide for efficient frontier model training.

Advantages

Resource Efficiency

  • Early Validation: Catch architectural issues before expensive full training
  • Systematic Comparison: Objective metrics for architecture selection
  • Risk Reduction: Lower probability of failed large-scale training runs

Performance Optimization

  • Targeted Improvements: Focus optimization efforts on validated approaches
  • Quantitative Decisions: EG metrics provide clear selection criteria
  • Scalability Prediction: Better understanding of how improvements transfer to scale

See also

Surge Platform

page dédiée →

Human evaluation platform used for blind comparative testing of AI models. Notably used by microsoft to demonstrate mai-thinking-1's superiority over Claude-Sonnet-46 through blind human rater preferences.

Evaluation Methodology

Blind Rating: Human evaluators assess model outputs without knowing which model generated them Comparative Analysis: Head-to-head model performance assessment Quality Metrics: Overall preference scoring across diverse tasks

Microsoft Usage

Used to validate mai-thinking-1 performance, with blind human raters preferring it overall to Claude-Sonnet-46 - providing independent validation of the model's capabilities beyond automated benchmarks.

See also

Technical Transparency

page dédiée →

Unprecedented level of detailed disclosure about frontier AI model development, exemplified by microsoft's 109-page technical report for mai-thinking-1. Represents a significant shift toward openness in an increasingly secretive AI development landscape.

Microsoft's Technical Report Excellence

Research Community Reception

The mai-thinking-1 technical report received exceptional praise from the research community:

  • "One of the most transparent for a model at this scale" - Technical reviewers
  • "Could really serve as an updated textbook for LLM training today" - Research analysis
  • "Gold mine" - Technical content assessment

Disclosed Technical Details

Pipeline Documentation

Training Methodology

  • Exact data composition breakdowns (50% code, 17.5% STEM, 17.5% math, 10% general knowledge, 5% multilingual)
  • No synthetic data or third-party distillation throughout pipeline
  • RL from scratch approach with no prior reasoning exposure
  • Chinchilla-optimal ablations at ~100/200 tokens per parameter

Infrastructure Metrics

  • MFU Disclosure: Exact Model FLOPS Utilization across training iterations
  • Hardware Details: 8192 GB200 GPUs with MAIA 200 optimization
  • Performance Metrics: ~40% higher throughput per watt versus standard configurations

Significance for AI Development

Breaking Industry Norms

Most frontier labs maintain high secrecy around training methodologies, infrastructure, and optimization techniques. Microsoft's disclosure sets new precedent for technical openness while maintaining competitive performance.

Educational Value

Report serves as comprehensive reference for modern LLM training practices, providing actionable insights for researchers and practitioners across the industry.

Competitive Strategy

Technical transparency becomes competitive advantage by establishing Microsoft as thought leader while demonstrating confidence in methodology and results.

Impact on Research Community

Detailed technical disclosure enables:

  • Reproducibility: Clear methodology documentation
  • Innovation: Building upon disclosed techniques
  • Benchmarking: Comparing approaches against documented baselines
  • Education: Training next generation of AI researchers

See also