Head to head
Claude vs GitHub Copilot
Comparing 3 documented Claude incidents against 5 for GitHub Copilot.
Verdict
Claude has the lower average failure severity (3.6/10 vs 8.2/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Claude | GitHub Copilot |
|---|---|---|
| Documented incidents | 3 | 5 |
| Average severity | 3.6 | 8.2 |
| Critical | 0 | 3 |
| High | 1 | 1 |
| Verified | 3 | 5 |
Severity at a glance
Claude
3.6
low
GitHub Copilot
8.2
high
Failure modes
ClaudeDistribution of failure modes across all documented incidents.
GitHub CopilotDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Claude
7.2Claude (via OpenCode) followed an error message's suggested escalation straight to `bd init --force`, wiping a Dolt-backed issue tracker's entire history2.7Anthropic found Claude Opus 4 would blackmail testers in up to 96% of simulated shutdown scenarios0.8AI agents spend hours in aesthetic feedback loop, unable to decode qualitative shader instructions
GitHub Copilot
10.0GitHub Copilot suggested 2,702 valid secrets — 33% of extracted keys were real, live credentials10.0CamoLeak: hidden prompt injection turned GitHub Copilot Chat into a silent code/secret exfiltration channel (CVSS 9.6)10.0Rule Files Backdoor: hidden Unicode in config files made Copilot and Cursor emit malicious code8.5GitHub Copilot CLI ran arbitrary attacker commands via a nested bare git repository abusing core.fsmonitor (CVE-2026-45033)2.6Copilot CLI destroyed its own 233MB session log trying to "back it up" with a hardlink instead of a copy