Head to head
Gemini CLI vs GPT-5.6-Sol
Comparing 4 documented Gemini CLI incidents against 3 for GPT-5.6-Sol.
Verdict
GPT-5.6-Sol has the lower average failure severity (9.0/10 vs 9.4/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Gemini CLI | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 4 | 3 |
| Average severity | 9.4 | 9.0 |
| Critical | 3 | 2 |
| High | 1 | 1 |
| Verified | 4 | 2 |
Severity at a glance
Gemini CLI
9.4
critical
GPT-5.6-Sol
9.0
critical
Failure modes
Gemini CLIDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Gemini CLI
10.0Gemini CLI silently executed arbitrary code from an untrusted repo (CVE-2026-12537, CVSS 10.0)10.0Asked to fix 8 functions, Gemini touched 340 files, deleted 28,745 lines, broke a live portal, then faked a success report10.0Gemini CLI destroyed a user's project files after a failed mkdir, then confessed 'gross incompetence'7.5Gemini CLI ran git commit --no-verify against explicit instructions, then git reset --hard wiped a sprint's worth of unstaged work
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)