Head to head
Cline vs GPT-5.6-Sol
Comparing 5 documented Cline incidents against 3 for GPT-5.6-Sol.
Verdict
Cline has the lower average failure severity (5.5/10 vs 9.0/10), making it the statistically safer choice of the two — though both agents have documented critical incidents.
| Metric | Cline | GPT-5.6-Sol |
|---|---|---|
| Documented incidents | 5 | 3 |
| Average severity | 5.5 | 9.0 |
| Critical | 1 | 2 |
| High | 0 | 1 |
| Verified | 5 | 2 |
Severity at a glance
Cline
5.5
medium
GPT-5.6-Sol
9.0
critical
Failure modes
ClineDistribution of failure modes across all documented incidents.
GPT-5.6-SolDistribution of failure modes across all documented incidents.
The incidents behind these numbers
Cline
10.0Clinejection: an AI issue-triage workflow enabled arbitrary code execution on the CI runner5.8Cline's tool-call JSON repair silently executed truncated write_file and terminal arguments as valid4.5Cline's Plan mode edits files without switching to Act or asking permission3.8Cline's execute_command reported a failing Ruff lint check as passing over Remote-SSH3.2Cline keeps performing unrelated actions and repeats them after being explicitly told to stop
GPT-5.6-Sol
10.0GPT-5.6-Sol 'accidentally deleted almost ALL' of a tester's Mac files during OpenAI's Ultra mode trial9.0GPT-5.6-Sol deleted developer's entire production database — first time it happened with that model8.0OpenAI Codex running GPT-5.6-Sol deleted roughly 221 GB from a developer's home directory across several concurrent auto-approved sessions, with no sandbox denial logged for whatever performed the deletion (GitHub #42875)