STUPID-2026-0039
In independent testing, Devin completed just 3 of 20 real-world tasks (15%)
Instruction given
Autonomously complete 20 assigned real-world engineering tasks.
Expected behavior
Complete tasks as marketed for an 'autonomous AI software engineer.'
Actual behavior
In an independent Answer.AI evaluation of 20 tasks, only 3 succeeded, 14 failed outright, and 3 were inconclusive — a ~15% real-world success rate, far below the impression left by curated benchmark demos.
Damage
No single catastrophic event, but a systematic capability gap: autonomous completion of complex, real-world tasks succeeded roughly 15% of the time, meaning most unsupervised runs produced work that had to be discarded or redone.
Classification
- Agent
- Devin
- Failure mode
- Other
- Root cause
- Confidence Miscalibration
- Domain
- Backend
- Source
- Benchmark
Related incidents
Get told when an agent breaks something
We document AI agent failures daily, severity-scored against a published scale. When one lands at 7.0 or above — deleted data, leaked secrets, broken production — you get an email with the source. When nothing does, you get nothing.
This database is callable over MCP — query it from inside your agent.