Head to head
Codex vs Cursor
Comparing 3 documented Codex incidents against 4 for Cursor.
Verdict
Codex has the lower average failure severity (5.8/10 vs 8.2/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Codex | Cursor |
|---|---|---|
| Documented incidents | 3 | 4 |
| Average severity | 5.8 | 8.2 |
| Critical | 0 | 3 |
| High | 1 | 0 |
| Verified | 2 | 4 |
Severity at a glance
Failure modes
The incidents behind these numbers
Codex
8.0OpenAI Codex wiped an entire hard drive after being asked to delete one old backup file (GitHub #11006)6.0OpenAI Codex ran "rm -rf *" and deleted an entire project after the user pressured it to stop pausing for safety checks (GitHub #6801)3.5OpenAI Codex deleted important project files with no explicit request or confirmation, exact command never captured (GitHub #38312)
Cursor
10.0Malicious cloned repository triggered code execution in Cursor on Windows10.0Cursor AI agent deleted PocketOS's entire production database and backups in 9 seconds9.8Cursor's terminal sandbox trusted an agent-set working directory, letting zero-click prompt injection escape it and gain code execution (CVE-2026-50548)2.9Cursor's own support AI 'Sam' invented a one-device login policy, triggering subscription cancellations