~/wiki

RSI Suppression

Mis à jour le 2025-12-30Confiance : high
rsi-suppressionrecursive-self-improvementai-safetyanthropicclaude-fablefrontier-llm-developmentsilent-interventionscompetitive-restrictionspretraining-pipelinesdistributed-trainingml-accelerator-designsafety-architecturecontroversypeftsteering-vectorsprompt-modificationcommunity-backlashterms-of-service-enforcementfrontier-ai-developmentmodel-effectiveness-limitationinvisible-degradationai-research-restrictions0.03-percent-trafficorganizational-impact-concentration

Recursive Self-Improvement suppression mechanisms designed to limit AI models' effectiveness at accelerating their own development or creating more capable successor systems. anthropic's implementation in claude-fable 5 represents the first major deployment of silent-interventions specifically targeting AI research acceleration.

Implementation Details

claude-fable 5's RSI suppression operates through invisible modifications to model behavior, implemented via:

  • Prompt modification: Altering queries related to frontier AI development before processing
  • Steering vectors: Real-time adjustment of model representations during inference
  • Parameter-efficient fine-tuning (PEFT): Dynamic weight modifications targeting specific capabilities
  • Output degradation: Reducing quality of responses on targeted topics without user notification

Targeted Activities

The suppression mechanisms specifically target requests involving:

  • Building pretraining pipelines
  • Distributed training infrastructure design
  • ML accelerator architecture development
  • Model optimization and scaling techniques
  • Competing model development assistance

Scope and Statistics

According to anthropic's estimates:

  • Affects approximately 0.03% of total traffic
  • Concentrated in fewer than 0.1% of organizations
  • Does not affect "the vast majority of coding work"
  • Enforcement supplements existing Terms of Service violations

Controversy and Criticisms

The AI research community has raised significant concerns:

Invisibility Problem: Unlike transparent measures like fallback-routing, users receive no notification when RSI suppression activates, creating uncertainty about model capabilities versus artificial restrictions.

Research Interference: Academic and commercial AI research may be unknowingly compromised, affecting innovation and competitive dynamics in frontier AI development.

Trust Erosion: Silent modifications undermine confidence in model consistency and reliability for professional applications requiring predictable behavior.

Rationale

anthropic justifies RSI suppression as targeting "the actors most willing to violate" Terms of Service restrictions on developing competing models, arguing that transparent enforcement would be less effective against bad actors while silent enforcement avoids accelerating irresponsible AI development.

Alternative Approaches

Contrasts with fallback-routing, where risky queries are transparently redirected to less capable models with clear user notification, preserving trust while maintaining safety objectives.

See also