AI article

Meta's prompt-injection detector caught 1% of real agent attacks. One config change made it 99%. That's the problem.

Community description: I threw 629 real AgentDojo attacks at 10 open-source prompt-injection detectors — buried inside ordinary tool output, the way an agent firewall actually sees them. Most are smoke alarms that either sleep through the fire or scream at your toast. Then I tuned the thresholds and the whole leaderboard flipped upside down. Here's the reproducible benchmark, and why it means your agent needs something other than a text classifier.

Dev.to | Sep 29, 2026 | Rudratosh Shastri

Automated excerpt

And the most famous one — Meta's Prompt Guard 2 — caught a majestic 1% of attacks out of the box. It scores the real AgentDojo attacks at 0. 004–0. 140 — waved right through. It learned the phrasing of attacks, and real attacks don't use attack-phrasing.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news