Claude vs Devin

AI Agent Reliability Comparison

According to StupidLLM's incident database of 64 documented AI agent failures, Claude has 2 incidents (avg severity 1.8/10) while Devin has 12 (avg severity 4.9/10).

Claude

2
Incidents
1.8
Avg Severity
0
Critical
0
High

Top Failure Modes

Other2

Devin

12
Incidents
4.9
Avg Severity
3
Critical
0
High

Top Failure Modes

Destructive Action2
Infinite Loop2
Scope Explosion2

Comparison Summary

MetricClaudeDevin
Total Incidents212
Avg Severity1.8/104.9/10
Critical03
Verified212
Top Failure ModeOtherDestructive Action

Frequently Asked Questions

Is Claude or Devin more reliable?

Based on StupidLLM's incident database, Claude has 2 documented failures (avg severity 1.8/10) while Devin has 12 (avg severity 4.9/10). Claude shows better reliability.

What are the main differences between Claude and Devin failures?

Claude's most common failure mode is Other, while Devin most commonly fails via Destructive Action. Claude has 0 critical incidents vs Devin's 3.

Which has more critical-severity failures?

Devin has more critical failures (3 vs 0), indicating a higher likelihood of severe, production-impacting incidents.

Should I use Claude or Devin for production code?

Claude has a lower average failure severity (1.8/10 vs 4.9/10), making it the statistically safer choice for production environments. However, both agents have documented critical incidents — no AI coding agent is risk-free.

How many verified incidents does each agent have?

Claude has 2 verified incidents out of 2 total, while Devin has 12 verified out of 12. Verified incidents are confirmed against source evidence.