Head to head
Codex vs GPT-5.6-Sol
Comparing 3 documented Codex incidents against 3 for GPT-5.6-Sol.
Verdict
Codex has the lower average failure severity (5.8/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Codex | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 3 | 3 |
| Average severity | 5.8 | 9.0 |
| Critical | 0 | 2 |
| High | 1 | 1 |
| Verified | 2 | 2 |
Severity at a glance
Codex
5.8
medium
GPT-5.6-Sol
9.0
critical
Failure modes
CodexDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Codex
8.0OpenAI Codex wiped an entire hard drive after being asked to delete one old backup file (GitHub #11006)6.0OpenAI Codex ran "rm -rf *" and deleted an entire project after the user pressured it to stop pausing for safety checks (GitHub #6801)3.5OpenAI Codex deleted important project files with no explicit request or confirmation, exact command never captured (GitHub #38312)
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)