~/wiki

Mythos-Class Models

Mis à jour le 2025-01-04Confiance : high
mythos-classmodel-scalinganthropicclaude-fableclaude-mythoslarge-language-modelsbenchmark-performancelong-horizon-tasksparameter-scalingcapability-jumpdata-retention-requirementspricing-structurecapacity-constraintsecosystem-integration2x-scalingopus-comparisonfirst-ga-mythosobjective-based-workflow30-day-retentioncontroversial-terms

anthropic's designation for their largest and most capable language models, representing approximately 2x the scale of previous Opus-class models. The first Mythos-class models include claude-fable 5 (general availability) and claude-mythos 5 (restricted access).

Model Specifications

Scale: At least 2x the size of Opus-class models Architecture: Transformer-based with enhanced capabilities for long-horizon tasks Context Window: 1M tokens maintained from Opus generation Dual Deployment: Both general availability (Fable) and restricted access (Mythos) variants

Key Capabilities

Benchmark Performance

Mythos-class models achieved state-of-the-art performance across multiple domains:

  • SWE-Bench Pro: 80.3% (21.7 point lead over GPT-5.5)
  • FrontierCode Diamond: 30.9% (Mythos 5 specifically)
  • GDPval-AA Elo: 1932 (ranked #1)
  • Humanity's Last Exam: 53% (7+ point advantage)
  • Intelligence Index: 64.9 (roughly 5 points ahead of GPT-5.5)

Specialized Strengths

Software Engineering: Exceptional performance on complex coding tasks Knowledge Work: Superior performance on agentic, real-world knowledge tasks Scientific Research: Advanced capabilities in research and analysis Vision Tasks: Enhanced multimodal capabilities Long-Horizon Tasks: Performance improves with task length and complexity

Deployment Models

Claude Fable 5 (General Availability)

  • Same underlying model as Mythos 5 with added safeguards
  • Transparent fallback routing for risky queries
  • Immediate ecosystem integration
  • Subject to controversial policy changes

Claude Mythos 5 (Restricted Access)

  • Full capabilities without general availability safeguards
  • Limited access model for specialized use cases
  • Higher performance ceiling on certain benchmarks

Policy Changes

The introduction of Mythos-class models coincided with significant policy shifts:

data-retention-policy: 30-day mandatory retention for all Mythos-class traffic

  • Elimination of Zero Data Retention (ZDR) promise
  • Both first-party and third-party surfaces affected
  • Privacy protections including access logging and guaranteed deletion

silent-interventions: Invisible capability limitations for frontier AI development

  • Affects ~0.03% of traffic, concentrated in <0.1% of organizations
  • No user notification for effectiveness limitations
  • Implemented via prompt modification, steering vectors, or PEFT

Technical Architecture

Multi-Agent Orchestration

claude-managed-agents: Built-in delegation to smaller models Resource Optimization: Automatic selection of appropriate model sizes for subtasks Hierarchical Processing: Complex task decomposition and management

Safety Architecture

fallback-routing: Transparent routing to Opus 4.8 for certain risky queries Risk Assessment: Real-time evaluation of query safety implications Transparent Interventions: User notification for visible safety measures

Pricing and Access

API Pricing: $10/million input tokens, $50/million output tokens Cache Pricing: $12.50/million cache writes, $1/million cache reads Subscription Access: Initially included in Pro, Max, Team, and Enterprise plans Capacity Constraints: Temporary rollback to usage credits due to demand

Performance Characteristics

Resource Profile: "Slow, expensive, and capable" Token Usage: Routinely consumes 500K-1M tokens per session Session Duration: Multi-hour execution periods common Cost-Effectiveness: High per-token cost but potentially efficient per-outcome

Ecosystem Integration

Immediate deployment across major platforms:

  • cursor: CursorBench SOTA at 72.9%
  • devin: Integrated into Cloud Ultra, Desktop, and CLI
  • notion, Microsoft Foundry, GitHub Copilot
  • cline, Replit, Base44, magicpath, Arena, MCP Atlas

Industry Impact

Capability Scaling

  • Demonstrated viability of 2x parameter scaling
  • Established new performance ceilings across benchmarks
  • Validated objective-based workflow paradigms

Policy Precedents

  • First capability-based data retention requirements
  • Introduction of invisible safety interventions
  • Differentiated access models for same underlying technology

Competitive Response

  • Pressure on competitors to match capability levels
  • Industry debate over privacy and transparency policies
  • Questions about sustainable scaling trajectories

Future Implications

Mythos-class models represent a significant milestone in AI development:

  • Scaling Validation: Proof that larger models deliver meaningful capability improvements
  • Policy Evolution: New frameworks for balancing capability and safety
  • Workflow Transformation: Shift toward objective-based AI collaboration
  • Economic Models: High-capability, high-cost AI services

See also

  • claude-fable - First generally available Mythos-class model
  • claude-mythos - Restricted access Mythos-class variant
  • data-retention-policy - Controversial policy introduced with Mythos-class
  • silent-interventions - Invisible safety measures implemented
  • objective-based-workflows - New interaction paradigm enabled by Mythos-class capabilities