Methodology
StupidLLM uses a CVSS-inspired severity scoring system to rate every documented AI agent failure on a 0-10 scale. Each incident is verified against source evidence and categorized by failure mode and root cause.
Severity Scoring (0-10)
Causes irreversible damage: data loss, security breach, production outage.
Significant impact: corrupted codebase, API key exposure, system instability.
Moderate impact: logic errors, wasted resources, incorrect but recoverable output.
Minor issues: cosmetic bugs, harmless hallucinations, non-production impact.
Verification
Every incident includes a source URL (GitHub PR, tweet, blog post, news article) used to verify the claim. Verified incidents are confirmed against their source evidence. Unverified incidents are marked as such and may be updated when sources become available.