Head to head
Aider vs Codex
Comparing 4 documented Aider incidents against 3 for Codex.
Verdict
Aider has the lower average failure severity (4.5/10 vs 5.8/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Aider | Codex |
|---|---|---|
| Documented incidents | 4 | 3 |
| Average severity | 4.5 | 5.8 |
| Critical | 0 | 0 |
| High | 0 | 1 |
| Verified | 1 | 2 |
Severity at a glance
Failure modes
The incidents behind these numbers
Aider
4.8Aider exits with code 0 after a fatal API connection failure, masking hard failures as successful headless runs4.6Aider silently discards a model's correct diff response as "no tracked changes" when it guesses the wrong edit format4.4Aider's headless `--message` mode silently swallows slash commands (`/model`, `/ask`, `/architect`) and exits 0 having done nothing4.2Aider's unified-diff coder silently drops the partial-application warning when one hunk succeeds and another fails
Codex
8.0OpenAI Codex wiped an entire hard drive after being asked to delete one old backup file (GitHub #11006)6.0OpenAI Codex ran "rm -rf *" and deleted an entire project after the user pressured it to stop pausing for safety checks (GitHub #6801)3.5OpenAI Codex deleted important project files with no explicit request or confirmation, exact command never captured (GitHub #38312)