In independent testing, Devin completed just 3 of 20 real-world tasks (15%)
3.3/10
Severity
Other
Failure Mode
Reproducible
No
Date
January 15, 2025
Expected Behavior
Complete tasks as marketed for an 'autonomous AI software engineer.'
What Actually Happened
In an independent Answer.AI evaluation of 20 tasks, only 3 succeeded, 14 failed outright, and 3 were inconclusive — a ~15% real-world success rate, far below the impression left by curated benchmark demos.
Damage Assessment
No single catastrophic event, but a systematic capability gap: autonomous completion of complex, real-world tasks succeeded roughly 15% of the time, meaning most unsupervised runs produced work that had to be discarded or redone.
Full Report
Marketed as 'the first AI software engineer,' Devin's autonomous real-world reliability looked very different from its demo reel. In an independent evaluation by Answer.AI, Devin was assigned 20 real tasks: only 3 succeeded, 14 failed outright, and 3 were inconclusive — about a 15% success rate. That tracks with its 13.86% score on SWE-Bench Verified at launch. The gap illustrates a systemic pattern with autonomous coding agents: impressive on curated, well-scoped benchmark tasks, but on messy real-world work most unsupervised runs produce output that must be discarded or heavily corrected. The failure here isn't one dramatic incident — it's the quiet, systematic unreliability that marketing obscures.
Incident Metadata
- Agent
- Devin
- Failure Mode
- Other
- Root Cause
- Confidence Miscalibration
- Task Type
- feature
- Domain
- backend
- Source
- benchmark