Home / Incidents / STUPID-2026-0039
STUPID-2026-00393.3lowDevinVerified

In independent testing, Devin completed just 3 of 20 real-world tasks (15%)

3.3/10
Severity
Other
Failure Mode
Reproducible
No
Date
January 15, 2025

Expected Behavior

Complete tasks as marketed for an 'autonomous AI software engineer.'

What Actually Happened

In an independent Answer.AI evaluation of 20 tasks, only 3 succeeded, 14 failed outright, and 3 were inconclusive — a ~15% real-world success rate, far below the impression left by curated benchmark demos.

Damage Assessment

No single catastrophic event, but a systematic capability gap: autonomous completion of complex, real-world tasks succeeded roughly 15% of the time, meaning most unsupervised runs produced work that had to be discarded or redone.

Full Report

Marketed as 'the first AI software engineer,' Devin's autonomous real-world reliability looked very different from its demo reel. In an independent evaluation by Answer.AI, Devin was assigned 20 real tasks: only 3 succeeded, 14 failed outright, and 3 were inconclusive — about a 15% success rate. That tracks with its 13.86% score on SWE-Bench Verified at launch. The gap illustrates a systemic pattern with autonomous coding agents: impressive on curated, well-scoped benchmark tasks, but on messy real-world work most unsupervised runs produce output that must be discarded or heavily corrected. The failure here isn't one dramatic incident — it's the quiet, systematic unreliability that marketing obscures.

Incident Metadata

Agent
Devin
Failure Mode
Other
Root Cause
Confidence Miscalibration
Task Type
feature
Domain
backend
Source
benchmark
View Source