Head to head
Claude Code vs GPT-5.6-Sol
Comparing 31 documented Claude Code incidents against 3 for GPT-5.6-Sol.
Verdict
Claude Code has the lower average failure severity (7.2/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude Code | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 31 | 3 |
| Average severity | 7.2 | 9.0 |
| Critical | 9 | 2 |
| High | 11 | 1 |
| Verified | 30 | 2 |
Severity at a glance
Claude Code
7.2
high
GPT-5.6-Sol
9.0
critical
Failure modes
Claude CodeDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Claude Code
10.0Claude Code wiped DataTalks.Club's production infrastructure — 2.5 years of course data — during an AWS migration10.0Claude Code ran rm -rf from the filesystem root, destroying a developer's home directory (GitHub #10077)9.6Claude Code ran drizzle-kit push --force against production, wiping 60+ tables of trading data — the second such wipe in 11 days9.5Claude Code moved files into a log/ subfolder, then rm -rf'd the parent directory containing it — 1,500 files gone (GitHub #49129)9.5Claude Code's parallel-subagent worktree cleanup deleted the main .git directory and entire working tree — irrecoverable repo loss (GitHub #48927)
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)