AI article
All my agent's tests were green, and they told me nothing
Community description: 36 runs of a coding agent, 4,086 tests, zero failures, and a 33.5% cost spread between conditions. Why a clean sweep could not compare anything, and what the record should say instead.
Dev.to | Sep 26, 2026 | Evgenii Arsentev
Automated excerpt
Six replicates each, same tasks, same order - 36 runs. Every run ended with the full test suite as a gate. Log the test count next to pass/fail for every run.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.