Head to head
Claude vs Cursor
Comparing 3 documented Claude incidents against 4 for Cursor.
Verdict
Claude has the lower average failure severity (3.6/10 vs 8.2/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude | Cursor |
|---|---|---|
| Documented incidents | 3 | 4 |
| Average severity | 3.6 | 8.2 |
| Critical | 0 | 3 |
| High | 1 | 0 |
| Verified | 3 | 4 |
Severity at a glance
Failure modes
The incidents behind these numbers
Claude
7.2Claude (via OpenCode) followed an error message's suggested escalation straight to `bd init --force`, wiping a Dolt-backed issue tracker's entire history2.7Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios0.8AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
Cursor
10.0Malicious cloned repository triggered code execution in Cursor on Windows10.0Cursor AI agent deleted PocketOS's entire production database and backups in 9 seconds9.8Cursor's terminal sandbox trusted an agent-set working directory, letting zero-click prompt injection escape it and gain code execution (CVE-2026-50548)2.9Cursor's own support AI 'Sam' invented a one-device login policy, triggering subscription cancellations