Concepts — vue longue
retour à la listeToutes les pages concaténées sur un seul document, pour un Ctrl-F direct.
Baseten Partnership
page dédiée →Strategic partnership between microsoft and baseten providing enterprise-controlled fine-tuning for mai-models with "100% eyes-off" data privacy guarantees.
Key Features
100% Eyes-Off: Complete data privacy - no human access to enterprise training data Enterprise-Controlled Fine-tuning: Customer maintains full control over model customization clean-data-lineage: Maintained throughout the fine-tuning process Privacy Compliance: Addresses enterprise data governance requirements
Strategic Value
Enterprise Adoption: Removes data privacy barriers for large organization AI deployment Competitive Advantage: Differentiates MAI models from competitors on privacy grounds Market Positioning: Positions Microsoft as enterprise-first AI provider
Technical Implementation
Built on mai-thinking-1 and broader MAI model family, enabling organizations to create specialized versions while maintaining data confidentiality and regulatory compliance.
See also
- mai-models
- enterprise-ai
- Data-Privacy
- clean-data-lineage
Chinchilla Optimal
page dédiée →Training methodology that balances model parameters and training tokens to achieve compute-optimal performance, based on scaling law research. Referenced in mai-thinking-1 development as a baseline for ablation studies.
MAI-Thinking-1 Implementation
Microsoft conducted ablations at roughly 100-200 tokens per parameter, described as "around Chinchilla optimal" for their setup, though noting differences from dense-model heuristics due to moe-architecture structure.
Scaling Law Foundation
Based on research demonstrating optimal compute allocation between:
- Model parameter count
- Training token quantity
- Computational budget constraints
MoE Considerations
Traditional Chinchilla optimal ratios may require adjustment for Mixture of Experts architectures due to different parameter utilization patterns and effective model capacity calculations.
Strategic Implications
Chinchilla-optimal training enables:
- Maximum performance for given compute budget
- Efficient resource allocation decisions
- Comparative evaluation of architecture efficiency
- Foundation for scaling law extrapolation
Research Impact
Established fundamental principles for compute-efficient training that influence model development decisions across the industry, providing scientific basis for training resource allocation.
See also
- mai-thinking-1
- Scaling-Laws
- moe-architecture
- Compute-Optimal-Training
Clean Data Lineage
page dédiée →Training methodology emphasizing transparent, traceable data sources without third-party model distillation or synthetic data generation. Pioneered by microsoft in the mai-models family, particularly mai-thinking-1.
Core Principles
No Distillation: Zero use of outputs from third-party models during training
No Synthetic Data: Reliance on authentic, naturally-occurring data sources
Transparent Sources: Clear documentation of all data origins and processing steps
Quality Control: Rigorous extraction, deduplication, and curation processes
Microsoft's Implementation
Data Sources:
- common-crawl web data
- Private, curated datasets
- Targeted sub-pipelines for different domains
Quality Assurance:
- Heavy extraction and deduplication work
- DSPy-GEPA optimized LLM judges for quality scoring
- Domain-specific curation pipelines
Enterprise Value
Trust: Clear data provenance for compliance and auditing Control: No dependency on competitor model outputs Quality: Higher signal-to-noise ratio through careful curation Legal Safety: Reduced IP and licensing complications
Industry Impact
Represents pushback against widespread use of synthetic data and model distillation, emphasizing the value of authentic data sources for frontier model development.
See also
- mai-models
- technical-transparency
- DSPy-GEPA
- Data-Curation
DSPy
page dédiée →Framework for optimizing language model pipelines through systematic prompt engineering and quality assessment. Notably used by microsoft in mai-thinking-1 development for advanced data curation and quality scoring through optimized LLM judges.
Core Concepts
DSPy provides systematic approaches to:
- Prompt Optimization: Automatic improvement of prompts through optimization algorithms
- Pipeline Composition: Chaining multiple LLM operations with optimization
- Quality Assessment: Using optimized LLM judges for data scoring and curation
GEPA Integration
GEPA (Generalized Evaluation and Prompt Adaptation) represents an advanced DSPy application involving:
- Late-interaction techniques for efficient retrieval and scoring
- Optimized LLM judges for quality assessment
- Systematic prompt adaptation based on performance metrics
Microsoft MAI Implementation
In mai-thinking-1 development, Microsoft leveraged DSPy for:
Data Curation
- Quality Scoring: DSPy-optimized LLM judges evaluated training data quality
- Domain-Specific Pipelines: Targeted curation for different content domains
- Extraction and Deduplication: Systematic processing of Common Crawl and private sources
Training Pipeline Integration
- Pre-training Data: Quality scoring during data preparation
- Scaling Decisions: Optimization metrics for architecture promotion
- Efficiency Measurement: Systematic evaluation of training effectiveness
Research Community Impact
The disclosure of DSPy usage in MAI-Thinking-1 development generated significant attention from the late-interaction and optimization research communities, demonstrating practical applications of systematic prompt optimization at frontier model scale.
Technical Applications
LLM Judge Optimization
- Automatic improvement of evaluation prompts
- Consistency optimization across evaluation runs
- Domain-specific judge adaptation
Pipeline Orchestration
- Multi-stage processing with optimized transitions
- Error correction and quality gating
- Performance monitoring and adaptation
See also
- mai-thinking-1
- clean-data-lineage
- llm-evaluation-framework
- automated-evaluation
- quality-scoring
Efficiency Gain Metric
page dédiée →Metric used by microsoft in mai-thinking-1 development for making architecture promotion decisions during the scaling-ladder process. Measures how much extra compute the baseline architecture would need to match a candidate architecture's loss performance.
Application in MAI-Thinking-1
Architecture Comparison Framework
The Efficiency Gain metric provides quantitative basis for comparing different architectural choices by measuring their relative computational efficiency for achieving equivalent loss performance.
Promotion Decision Criteria
Used as primary metric for deciding which candidate architectures to promote to larger scales during the systematic scaling ladder evaluation process.
Compute Resource Optimization
Enables data-driven decisions about resource allocation by quantifying the computational cost difference between architectural alternatives.
Technical Implementation
Baseline vs Candidate Evaluation
- Baseline Architecture: Reference architecture with known compute requirements
- Candidate Architecture: Alternative architecture being evaluated
- Efficiency Calculation: Ratio of compute needed by baseline to match candidate's loss
Loss Performance Matching
The metric specifically measures compute requirements to achieve equivalent loss performance, rather than other metrics that might not directly correlate with training efficiency.
Scale-Aware Assessment
Applied across different compute scales during the scaling ladder process, ensuring architectural decisions remain optimal at various training scales.
Research Significance
Systematic Architecture Selection
Provides objective framework for architecture selection in large-scale model development, moving beyond intuitive or experience-based decisions.
Resource Planning
Enables predictive planning for computational resource requirements when scaling architectural decisions to full model training.
Reproducible Methodology
Creates standardized approach for architectural evaluation that can be applied systematically across different model development projects.
See also
- scaling-ladder
- mai-thinking-1
- Architecture-Optimization
- technical-transparency
GEPA
page dédiée →Advanced technique used in conjunction with dspy for optimizing LLM judges in data curation and quality scoring. Notably employed by microsoft in mai-thinking-1 development for pretraining data quality assessment.
Integration with DSPy
GEPA works within the DSPy framework to enhance LLM judge optimization, particularly for:
- Pretraining data curation
- Quality scoring of training examples
- Automated data pipeline evaluation
- Late-interaction optimization
MAI-Thinking-1 Implementation
Microsoft's use of DSPy-optimized LLM judges with GEPA represented a sophisticated approach to data quality control, contributing to the model's clean data lineage and high performance outcomes.
Technical Community Interest
Generated significant attention from the DSPy and late-interaction research communities, highlighting the growing importance of optimized evaluation systems in frontier model development.
Relationship to Data Quality
Part of Microsoft's broader emphasis on clean-data-lineage, demonstrating how advanced curation techniques can substitute for synthetic data or distillation approaches while maintaining high model performance.
See also
- dspy
- mai-thinking-1
- clean-data-lineage
- LLM-Judges
MAI Models
page dédiée →microsoft's internally developed language model series emphasizing clean lineage, exceptional data quality, and hill climbing capabilities. Designed to enable companies to build their own specialist models rather than relying solely on generalist models. At Build 2026, Microsoft announced seven new MAI models demonstrating competitive frontier capabilities and unprecedented technical-transparency.
Design Philosophy
Clean Lineage Foundation
Starting with pre-training using very high data quality with extensive ablation studies. satya-nadella emphasized this is "becoming even harder to build a clean lineage model just because there's so much stuff out there that you truly need to ablate out to be able to have a fantastic pre-trained model."
This addresses a key limitation of many open weight models that "look great on one benchmark or two, but they're not great on practice."
Cognitive Core Pursuit
Central to MAI development is pursuing the "cognitive-core" - fundamental intelligence patterns that can serve as the foundation for specialized capabilities. This approach prioritizes essential intelligence over pure scale.
Hill Climbing Architecture
Scaffold System
MAI models include a "hill climb scaffold" enabling customers to:
- Build specialist models from the generalist foundation
- Implement trace-collection for continuous improvement
- Develop private-evals specific to their domain
- Create proprietary intellectual property through model specialization
Temporal Scaffolding Innovation
Demonstrated through the land-o-lakes-demo where:
- GPT-55 was used to collect traces
- A 5B reasoning model achieved higher performance using those traces
- This represents "a new frontier" in AI capability development
Platform Integration Strategy
MAI models serve as the foundation for Microsoft's frontier-intelligence-platform approach:
- Enable "first-class participants" who can point to AI they created
- Support enterprise specialization rather than generic AI consumption
- Integrate with multi-model harnesses like openclaw and scout
- Connect with enterprise context through work-iq
Seven Model Family (Build 2026)
Microsoft announced seven new MAI models demonstrating:
- Competitive frontier capabilities
- Unprecedented technical transparency
- Specialized capabilities across different domains
- Support for enterprise-controlled fine-tuning
Training Strategy Advantages
Data Quality Focus
- Extensive ablation studies to ensure clean training data
- Careful curation to avoid contamination common in open models
- Focus on quality over quantity in training corpus
Specialized Development Path
- Not just generalist models but foundation for specialization
- Enables customers to develop proprietary AI capabilities
- Supports enterprise-specific use cases and requirements
Competitive Positioning
MAI models position Microsoft uniquely as:
- Both platform provider and frontier model developer
- Enabling customer AI development rather than just AI consumption
- Balancing technical capability with ecosystem enablement
- Addressing practical deployment challenges through clean architecture
See also
MAIA-200
page dédiée →microsoft's custom AI chip optimized for running mai-models, delivering significant performance and efficiency improvements over standard GB200-GPUs for Microsoft's model inference workloads.
Performance Characteristics
Cost Efficiency: 30% better performance per dollar compared to GB200 Power Efficiency: 1.4x performance-per-watt gain versus GB200 Optimization Target: End-to-end MAI model serving and inference
Strategic Importance
Hardware-Software Co-design: Custom silicon optimized specifically for MAI model architectures Cost Advantage: Significant operational cost benefits for Microsoft's AI services Competitive Moat: Hardware optimization as differentiation strategy
Technical Integration
Optimized for mai-thinking-1 and broader MAI family serving, representing Microsoft's investment in full-stack AI infrastructure control from silicon to software.
See also
- mai-models
- mai-thinking-1
- microsoft
- GB200-GPUs
MFU Disclosure
page dédiée →Model FLOPs Utilization (MFU) metrics disclosure representing the percentage of theoretical hardware performance achieved during training. microsoft's disclosure of exact MFU numbers across iterations for mai-thinking-1 was noted as unprecedented transparency for frontier model development.
Significance of Disclosure
MFU numbers are rarely shared at frontier model scale because they reveal:
- Infrastructure efficiency and capabilities
- Engineering quality and optimization expertise
- Competitive training cost information
- Hardware utilization optimization techniques
Technical Importance
MFU measurements enable:
- Objective comparison of training infrastructure efficiency
- Identification of optimization opportunities
- Hardware procurement and scaling decisions
- Engineering team performance assessment
Microsoft's Transparency
The disclosure of exact MFU across training iterations demonstrated technical-transparency that multiple researchers highlighted as "rarely shared at this scale," contributing to the research community's positive reception of the technical report.
Industry Impact
Such detailed disclosure sets new standards for frontier model transparency and provides valuable reference points for the broader AI research community working on training efficiency optimization.
See also
- mai-thinking-1
- technical-transparency
- Training-Efficiency
- Infrastructure-Optimization
MoE Architecture
page dédiée →Mixture of Experts (MoE) is a neural network architecture that uses multiple specialized sub-networks (experts) with a gating mechanism to route inputs to the most relevant experts. Enables efficient scaling by activating only a subset of total parameters for each inference, as demonstrated in mai-thinking-1 and other frontier models.
Core Concepts
Expert Specialization:
- Multiple specialized sub-networks within single model
- Each expert develops domain-specific capabilities
- Gating network learns to route inputs to appropriate experts
- Enables model specialization without full parameter activation
Parameter Efficiency:
- Total Parameters: Full model size including all experts
- Active Parameters: Subset activated for any given input
- Example: mai-thinking-1 has 1T total parameters but only 35B active
- Significant compute savings during inference while maintaining model capacity
Scaling Advantages
Compute Efficiency:
- Linear scaling of experts with sub-linear compute growth
- Better performance per active parameter compared to dense models
- Different scaling laws compared to traditional dense architectures
- More efficient than equivalent dense models at same active parameter count
Specialization Benefits:
- Experts can develop domain-specific knowledge (code, math, language-specific)
- Reduced interference between different capability areas
- Better performance on diverse tasks within single model
- Easier to add new capabilities through additional experts
Implementation Challenges
Training Complexity:
- Load balancing across experts to prevent expert collapse
- Routing efficiency and stability during training
- More complex infrastructure requirements
- Different scaling heuristics compared to dense models
Architecture Design:
- Optimal expert count and size determination
- Gating mechanism design and training
- Expert specialization encouragement vs generalization
- Memory and communication overhead management
Microsoft MAI Implementation
MAI-Thinking-1 Specifications:
- 1T total parameters with 35B active parameters
- ~28x parameter efficiency ratio
- 256K context window maintained across all experts
- Optimized for MAIA 200 custom hardware
MAI-Code-1-Flash:
- 137B total parameters with 5B active parameters
- ~27x parameter efficiency for coding tasks
- Achieves 51% SWE-Bench Pro despite small active footprint
- Specialized for VS Code and GitHub Copilot integration
Training Methodology
Scaling Decisions:
- Efficiency Gain (EG) metrics for architecture promotion
- Ablations around Chinchilla-optimal token ratios adapted for MoE
- Private evaluation sets for expert specialization assessment
- Custom scaling laws for MoE vs dense model comparison
Expert Development:
- Domain-specific data routing during training
- Load balancing mechanisms to ensure expert utilization
- Specialization encouragement through routing policies
- Quality control across expert outputs
Hardware Optimization
Custom Silicon Integration:
- MAIA 200 optimization providing 30% better performance per dollar
- 1.4x performance-per-watt gain for MAI models
- Hardware-software co-design for expert routing efficiency
- Memory hierarchy optimization for sparse activation patterns
Industry Context
Competitive Landscape:
- Most frontier models now use some form of MoE architecture
- Enables larger models without proportional compute increase
- Critical for cost-effective serving of large language models
- Standard approach for balancing capability with efficiency
Future Directions:
- Increasing expert counts and specialization
- Dynamic expert creation and pruning
- Cross-modal expert architectures
- More sophisticated routing mechanisms
MoE architecture represents a fundamental shift toward sparse, efficient scaling in language models, enabling the parameter counts necessary for frontier performance while maintaining practical deployment constraints.
See also
- mai-thinking-1
- mai-models
- parameter-efficiency
- model-scaling
- expert-specialization
Private NLL Evaluation
page dédiée →Internal evaluation methodology using Negative Log Likelihood (NLL) on private datasets for making scaling and architecture decisions during model development. Employed by microsoft in mai-thinking-1 development for systematic model progression.
Data Composition for MAI-Thinking-1
Microsoft's private NLL evaluation set comprised:
- 50% code
- 17.5% STEM
- 17.5% math
- 10% general knowledge
- 5% multilingual
Role in Scaling Decisions
Used to evaluate candidate architectures during the scaling-ladder process, providing consistent performance measurement across different model scales and configurations.
Technical Implementation
Negative Log Likelihood provides a fundamental loss measurement that enables:
- Objective comparison between model architectures
- Scaling law analysis and extrapolation
- Data-driven decisions on architecture promotion
- Consistent evaluation across training iterations
Strategic Advantage
Private evaluation sets enable companies to make scaling decisions based on proprietary benchmarks that may better reflect target use cases than public benchmarks, while maintaining evaluation consistency across development cycles.
See also
- mai-thinking-1
- scaling-ladder
- efficiency-gain-metric
- Negative-Log-Likelihood
Scaling Ladder
page dédiée →Training methodology used by microsoft in developing mai-thinking-1, involving systematic architecture evaluation and promotion decisions based on performance metrics across different compute scales.
Core Methodology
Architecture Evaluation Process
The scaling ladder involves testing candidate architectures at smaller scales before committing to full-scale training, allowing for data-driven decisions about which architectures to promote to larger scales.
Efficiency Gain Metric
Architecture promotion decisions are based on the efficiency-gain-metric, which quantifies how much extra compute the baseline architecture would need to match a candidate architecture's loss performance.
Systematic Scaling Decisions
Rather than intuitive or heuristic-based scaling choices, the methodology provides quantitative framework for architecture selection at each scale tier.
Application in MAI-Thinking-1
Ablation Studies
Conducted ablations at approximately 100/200 tokens per parameter, described as "Chinchilla optimal" for the MoE setup, though differing from dense model heuristics.
Data-Driven Promotion
Architecture candidates systematically evaluated and promoted based on performance metrics rather than subjective assessment or industry conventions.
Scale-Aware Optimization
Methodology accounts for the fact that optimal architectures may differ at various compute scales, particularly for MoE configurations.
Technical Innovation
Beyond Dense Model Heuristics
The scaling ladder methodology explicitly accounts for MoE architectural differences, recognizing that traditional dense model scaling laws may not apply directly.
Systematic Experimentation
Provides framework for rigorous experimental methodology in large-scale model development, moving beyond ad-hoc scaling decisions.
Resource Optimization
Enables efficient use of computational resources by making informed decisions about architecture promotion rather than training all candidates to full scale.
Research Community Impact
The detailed disclosure of scaling ladder methodology in Microsoft's technical report provides actionable framework for other researchers developing large-scale models with systematic architecture evaluation.
See also
- mai-thinking-1
- efficiency-gain-metric
- moe-architecture
- technical-transparency
Scaling Ladder Methodology
page dédiée →Systematic approach to model development that uses incremental scaling and architecture comparison to optimize model design before full-scale training. Prominently featured in microsoft's mai-thinking-1 development process.
Core Methodology
Efficiency Gain (EG) Metric
Architecture promotion decisions based on Efficiency Gain calculation:
- Definition: How much extra compute the baseline would need to match the candidate's loss
- Application: Systematic comparison of architectural variants
- Decision Framework: Objective criteria for architecture selection
Scaling Progression
Incremental scaling through defined checkpoints:
- Small Scale Testing: Initial architecture validation
- Progressive Scaling: Systematic increase in model size and training data
- Architecture Refinement: Continuous optimization based on EG metrics
Training Schedule Optimization
Ablation Studies
Conducted at approximately 100-200 tokens per parameter:
- Chinchilla Optimal Range: Roughly optimal for the experimental setup
- MoE Adaptations: Different from dense model heuristics due to moe-architecture
- Resource Efficiency: Systematic testing without full-scale resource commitment
Validation Methodology
- Internal NLL Set: Private validation dataset for scaling decisions
- Loss Tracking: Systematic monitoring of training loss across scales
- Performance Prediction: Extrapolation from smaller scale results
MAI-Thinking-1 Implementation
Data Composition
Internal validation set composition for scaling decisions:
- 50% code
- 17.5% STEM
- 17.5% math
- 10% general knowledge
- 5% multilingual
Architecture Decisions
- MoE Configuration: Optimal expert count and routing strategies
- Parameter Allocation: Balance between active and total parameters
- Context Window: Optimization for 256K token context length
Research Impact
The detailed disclosure of scaling ladder methodology in MAI-Thinking-1's technical report provides unprecedented insights into systematic model development, serving as a practical guide for efficient frontier model training.
Advantages
Resource Efficiency
- Early Validation: Catch architectural issues before expensive full training
- Systematic Comparison: Objective metrics for architecture selection
- Risk Reduction: Lower probability of failed large-scale training runs
Performance Optimization
- Targeted Improvements: Focus optimization efforts on validated approaches
- Quantitative Decisions: EG metrics provide clear selection criteria
- Scalability Prediction: Better understanding of how improvements transfer to scale
See also
- mai-thinking-1
- moe-architecture
- model-optimization
- ablation-studies
- training-optimization
- architecture-selection
Surge Platform
page dédiée →Human evaluation platform used for blind comparative testing of AI models. Notably used by microsoft to demonstrate mai-thinking-1's superiority over Claude-Sonnet-46 through blind human rater preferences.
Evaluation Methodology
Blind Rating: Human evaluators assess model outputs without knowing which model generated them Comparative Analysis: Head-to-head model performance assessment Quality Metrics: Overall preference scoring across diverse tasks
Microsoft Usage
Used to validate mai-thinking-1 performance, with blind human raters preferring it overall to Claude-Sonnet-46 - providing independent validation of the model's capabilities beyond automated benchmarks.
See also
- mai-thinking-1
- human-evaluation
- model-benchmarking
Technical Transparency
page dédiée →Unprecedented level of detailed disclosure about frontier AI model development, exemplified by microsoft's 109-page technical report for mai-thinking-1. Represents a significant shift toward openness in an increasingly secretive AI development landscape.
Microsoft's Technical Report Excellence
Research Community Reception
The mai-thinking-1 technical report received exceptional praise from the research community:
- "One of the most transparent for a model at this scale" - Technical reviewers
- "Could really serve as an updated textbook for LLM training today" - Research analysis
- "Gold mine" - Technical content assessment
Disclosed Technical Details
Pipeline Documentation
- Complete scaling-ladder methodology
- efficiency-gain-metric for architecture decisions
- Data curation processes using dspy GEPA-optimized judges
- Infrastructure metrics and hardware utilization
Training Methodology
- Exact data composition breakdowns (50% code, 17.5% STEM, 17.5% math, 10% general knowledge, 5% multilingual)
- No synthetic data or third-party distillation throughout pipeline
- RL from scratch approach with no prior reasoning exposure
- Chinchilla-optimal ablations at ~100/200 tokens per parameter
Infrastructure Metrics
- MFU Disclosure: Exact Model FLOPS Utilization across training iterations
- Hardware Details: 8192 GB200 GPUs with MAIA 200 optimization
- Performance Metrics: ~40% higher throughput per watt versus standard configurations
Significance for AI Development
Breaking Industry Norms
Most frontier labs maintain high secrecy around training methodologies, infrastructure, and optimization techniques. Microsoft's disclosure sets new precedent for technical openness while maintaining competitive performance.
Educational Value
Report serves as comprehensive reference for modern LLM training practices, providing actionable insights for researchers and practitioners across the industry.
Competitive Strategy
Technical transparency becomes competitive advantage by establishing Microsoft as thought leader while demonstrating confidence in methodology and results.
Impact on Research Community
Detailed technical disclosure enables:
- Reproducibility: Clear methodology documentation
- Innovation: Building upon disclosed techniques
- Benchmarking: Comparing approaches against documented baselines
- Education: Training next generation of AI researchers
See also
- mai-thinking-1
- microsoft
- scaling-ladder
- Research-Disclosure