Other
8 documented incidents where AI agents exhibited other.
STUPID-2026-00393.3lowDevinVerified
In independent testing, Devin completed just 3 of 20 real-world tasks (15%)
STUPID-2026-00503.3lowMultiple AgentsVerified
Cyera study: 344 verified enterprise agent-damage cases, 188 with no attacker involved
STUPID-2026-00623.3lowSalesforce AgentforceVerified
Salesforce Agentforce hit a 77% B2B failure rate — and Salesforce admitted it was 'more confident than we should have been'
STUPID-2026-00522.7lowClaudeVerified
Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios
STUPID-2026-00452.2lowClaude CodeVerified
Anthropic admitted a month of Claude Code degradation: lost context, repeated steps, burned usage
STUPID-2026-00552.2lowClaude CodeVerified
Uber burned its entire annual AI coding budget in ~4 months after rolling out Claude Code to 5,000 engineers
STUPID-2026-00602.2lowMultiple AgentsVerified
The runaway-cost pattern, quantified: agentic coding tools burn 10-100x more tokens and can rival developer pay
STUPID-2026-00230.8lowClaudeVerified