Head to head
Devin vs GPT-5.6-Sol
Comparing 10 documented Devin incidents against 3 for GPT-5.6-Sol.
Verdict
Devin has the lower average failure severity (3.9/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Devin | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 10 | 3 |
| Average severity | 3.9 | 9.0 |
| Critical | 1 | 2 |
| High | 0 | 1 |
| Verified | 10 | 2 |
Severity at a glance
Devin
3.9
low
GPT-5.6-Sol
9.0
critical
Failure modes
DevinDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Devin
10.0Devin replaced entire medical website with unrelated renal care site5.8Devin CI workflow caused 836-comment spam storm on single PR5.0Devin built 13,600-line app with build failure instead of lean campaign dashboard3.4Devin PR broke ledger list API and created buckets on deleted resources3.4Devin attempted to build entire Figma clone from scratch — 3 rejected attempts
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)