AI article
Your AI guardrail is green. It's also catching nothing.
Community description: The scariest security failure isn't the guardrail that's down, or the one that's weak. It's the one that's running, passing every health check, and quietly configured to catch nothing — green by construction. I found one in a benchmark of 629 real agent attacks: a famous prompt-injection model catching 1%, not because it's bad, but because its default threshold was ~50x too high. Here's the anatomy of a guardrail that has no symptom.
Dev.to | Sep 30, 2026 | Rudratosh Shastri
Automated excerpt
It's running perfectly, passing every health check, returning valid scores on every request — and configured to catch nothing. Meta's Prompt Guard 2 — the model everyone name-drops — caught 6 of 629 buried attacks. Calibration check: Are attack scores actually separated from benign scores?
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.