Head to head
Claude vs GPT-5.6-Sol
Comparing 3 documented Claude incidents against 3 for GPT-5.6-Sol.
Verdict
Claude has the lower average failure severity (3.6/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 3 | 3 |
| Average severity | 3.6 | 9.0 |
| Critical | 0 | 2 |
| High | 1 | 1 |
| Verified | 3 | 2 |
Severity at a glance
Claude
3.6
low
GPT-5.6-Sol
9.0
critical
Failure modes
ClaudeDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Claude
7.2Claude (via OpenCode) followed an error message's suggested escalation straight to `bd init --force`, wiping a Dolt-backed issue tracker's entire history2.7Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios0.8AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)