About
What StupidLLM is
When an AI coding agent deletes a directory, leaks a secret, or rewrites half a codebase it was never asked to touch, that failure usually lives in a single tweet or bug report and then disappears. StupidLLM assigns those failures a stable identifier, scores them against a published rubric, and makes them searchable.
The database currently holds 83 documented incidents across 24 agents, plus 14 further multi-agent or unattributed reports — 97 in total. Every entry records the instruction given, the expected behavior, what actually happened, and the resulting damage — plus a source link so the claim can be checked.
What this is not
It is not a benchmark. Incident counts reflect public scrutiny as much as reliability, and there is no denominator of total tasks attempted. The methodology page states the limitations in full. Read the rankings as documented evidence, not as a defect rate.
License
The incident data is published under CC BY 4.0. You may reuse it, including commercially, with attribution.
How to cite
Contributing
Corrections and new incidents are welcome — see report an incident. If you believe an entry is wrong, say so and it will be corrected or withdrawn.