StupidLLM

The incident database for AI agent failures

According to StupidLLM's incident database, 64 AI agent failures have been documented across 27 agents like Devin, Cursor, Claude Code, GitHub Copilot, Windsurf, and Aider — with an average severity of 6.7/10. Like CVE for cybersecurity vulnerabilities, we assign STUPID-IDs to documented cases where AI agents cause real damage — deleted files, security vulnerabilities, infinite loops, wasted resources, and broken production systems.

64
Incidents Documented
6.7
Avg Severity /10
27
Agents Tracked
64
Verified

Latest Incidents

Highest Severity

Most Tracked Agents

What is StupidLLM?

According to StupidLLM's incident database, 64 AI agent failures have been documented across 27 agents with an average severity of 6.7/10. Every incident is severity-scored using our CVSS-inspired rating system, verified against source evidence, and searchable by agent, failure mode, and root cause. We track reliability trends across agents so developers and enterprises can make informed decisions about which AI tools to trust.

How are AI agent incidents scored?

Every incident is severity-scored on a 0-10 scale using a CVSS-inspired rating system. Scores of 9-10 are critical, 7-8 are high, 4-6 are medium, and 0-3 are low severity. Incidents are verified against source evidence and categorized by failure mode and root cause.

Which AI coding agent has the most failures?

Visit the StupidLLM dashboard for live rankings of AI agent failure rates. We track 64 incidents across 27 agents with average severity scores and risk levels.

AI Agent Failure Modes