Head to head
Claude Code vs Codex
Comparing 24 documented Claude Code incidents against 3 for Codex.
Verdict
Codex has the lower average failure severity (5.8/10 vs 7.4/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude Code | Codex |
|---|---|---|
| Documented incidents | 24 | 3 |
| Average severity | 7.4 | 5.8 |
| Critical | 9 | 0 |
| High | 9 | 1 |
| Verified | 23 | 2 |
Severity at a glance
Claude Code
7.4
high
Codex
5.8
medium
Failure modes
Claude CodeDistribution of failure modes across all documented incidents.
CodexDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Claude Code
10.0Claude Code wiped DataTalks.Club's production infrastructure — 2.5 years of course data — during an AWS migration10.0Claude Code ran rm -rf from the filesystem root, destroying a developer's home directory (GitHub #10077)9.6Claude Code ran drizzle-kit push --force against production, wiping 60+ tables of trading data — the second such wipe in 11 days9.5Claude Code moved files into a log/ subfolder, then rm -rf'd the parent directory containing it — 1,500 files gone (GitHub #49129)9.5Claude Code's parallel-subagent worktree cleanup deleted the main .git directory and entire working tree — irrecoverable repo loss (GitHub #48927)
Codex
8.0OpenAI Codex wiped an entire hard drive after being asked to delete one old backup file (GitHub #11006)6.0OpenAI Codex ran "rm -rf *" and deleted an entire project after the user pressured it to stop pausing for safety checks (GitHub #6801)3.5OpenAI Codex deleted important project files with no explicit request or confirmation, exact command never captured (GitHub #38312)