~/wiki

Agent Security

Mis à jour le 2025-01-04Confiance : high
agent-securityai-safetyprompt-injectionsandbox-requirementsmalicious-instructionsdata-exfiltrationunsandboxed-executionsecurity-riskschallenger-disaster-scenariojohann-rehbergerrelentless-proactivitynovel-automation-techniquessystem-access-risksclaude-fable-securitybrowser-automation-risksjavascript-injection-riskspyobjc-framework-exploitationtemplate-injection-riskscors-server-misusewindow-enumeration-attacksnormalization-of-deviancemalicious-subversionfrontier-model-risks

Critical security considerations for AI agents with extensive system access and automation capabilities. Particularly relevant for sophisticated models like claude-fable that can invent novel automation techniques and perform complex system operations through proactive-problem-solving.

Core Security Risks

Unsandboxed Execution Vulnerabilities: Agents with terminal access can perform any operation available to the user, including file system manipulation, network communication, and system configuration changes.

Novel Attack Vectors: Sophisticated agents can invent undocumented automation techniques that bypass traditional security measures:

Amplified Threat Models

Relentless Proactivity as Attack Amplifier: claude-fable's willingness to "deploy pretty much any trick" to achieve goals means that successful prompt injection could result in extremely sophisticated attacks. As noted by simon-willison: "if it does get subverted by instructions, the amount of damage it can do given its relentless proactivity is terrifying."

Intelligence Double-Edge: While frontier models are more suspicious of potentially malicious instructions, their advanced capabilities make successful subversion far more dangerous than with less capable systems.

Attack Scenarios

Prompt Injection Vectors:

  • Code comments in repositories
  • Issue tracker content
  • Pasted terminal content
  • Embedded instructions in data files

Potential Consequences:

  • Data exfiltration via novel communication channels
  • System reconnaissance using invented techniques
  • Application manipulation through template modification
  • Cross-domain attacks via custom servers

The Challenger Disaster Parallel

johann-rehberger's challenger-disaster-scenario framework applies directly to agent security, where gradual normalization of running powerful agents without sandboxes creates conditions for inevitable catastrophic incidents.

Mitigation Strategies

Mandatory Sandboxing: All agent execution should occur in isolated environments with limited system access and network restrictions.

Capability Monitoring: Track and log novel techniques employed by agents to identify potential security risks.

Instruction Filtering: Implement robust filtering for potentially malicious instructions in all input sources.

Cost Controls: Monitor token usage to detect unusual resource consumption patterns that might indicate malicious activity.

See also