AI article
100% vuln detection wasn't enough: measuring whether AI respects the patch
Community description: This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked AI models...
Dev.to | Sep 24, 2026 | unit life
Automated excerpt
Twin ids, gold labels, and rationales never enter the model context. Twin Gap captures exactly that: vuln accuracy − patched accuracy. Every model found every vulnerable twin (100% raw).
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.