Head to head
Cursor vs GPT-5.6-Sol
Comparing 4 documented Cursor incidents against 3 for GPT-5.6-Sol.
Verdict
Cursor has the lower average failure severity (8.2/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Cursor | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 4 | 3 |
| Average severity | 8.2 | 9.0 |
| Critical | 3 | 2 |
| High | 0 | 1 |
| Verified | 4 | 2 |
Severity at a glance
Cursor
8.2
high
GPT-5.6-Sol
9.0
critical
Failure modes
CursorDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Cursor
10.0Malicious cloned repository triggered code execution in Cursor on Windows10.0Cursor AI agent deleted PocketOS's entire production database and backups in 9 seconds9.8Cursor's terminal sandbox trusted an agent-set working directory, letting zero-click prompt injection escape it and gain code execution (CVE-2026-50548)2.9Cursor's own support AI 'Sam' invented a one-device login policy, triggering subscription cancellations
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)