AI article

Semantic DLP: Can AI Models Catch the Data Leak?

Community description: This is a submission for the Kaggle Benchmarking Challenge What I Benchmarked Employees...

Dev.to | Oct 11, 2026 | karthik sethuraman

Automated excerpt

Catch rate tracks model class: gemini-2. 5-pro catches every buried leak (1. 000), flash misses 4% (0. 962), gpt-oss-120b 17% (0. 825), and the 21B gpt-oss-20b misses 1 in 4 (0. 750). The frontier model clears the accuracy bar; small open-weight reasoning models The recall champion is the false-alarm champion: gemini-2. 5-pro flags 39. 3% of clean hard negatives, worse than DeBERTa-v3's 36%, the local side's cautionary tale. The board extends via Add Models. thresholds; and extending the corpus to more leak categories.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

Read next

AI briefing: recent picks

More stories to explore