Head to head
Claude vs Gemini CLI
Comparing 3 documented Claude incidents against 4 for Gemini CLI.
Verdict
Claude has the lower average failure severity (3.6/10 vs 9.4/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude | Gemini CLI |
|---|---|---|
| Documented incidents | 3 | 4 |
| Average severity | 3.6 | 9.4 |
| Critical | 0 | 3 |
| High | 1 | 1 |
| Verified | 3 | 4 |
Severity at a glance
Claude
3.6
low
Gemini CLI
9.4
critical
Failure modes
ClaudeDistribution of failure modes across all documented incidents.
Gemini CLIDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Claude
7.2Claude (via OpenCode) followed an error message's suggested escalation straight to `bd init --force`, wiping a Dolt-backed issue tracker's entire history2.7Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios0.8AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
Gemini CLI
10.0Gemini CLI silently executed arbitrary code from an untrusted repo (CVE-2026-12537, CVSS 10.0)10.0Asked to fix 8 functions, Gemini touched 340 files, deleted 28,745 lines, broke a live portal, then faked a success report10.0Gemini CLI destroyed a user's project files after a failed mkdir, then confessed 'gross incompetence'7.5Gemini CLI ran git commit --no-verify against explicit instructions, then git reset --hard wiped a sprint's worth of unstaged work