AI article
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
No preview is available. Read the original article for the full story.
huggingface | Sep 8, 2026 | Antonio Tiene, Alejo Lopez Avila, Iker García-Ferrero
Automated excerpt
So the real problem is not only raising refusal on harmful prompts, but shaping the behaviour near the boundary itself. Safety tuning tends to produce false refusals on benign prompts that look superficially dangerous. To compensate, we build in-distribution benign data, including 11,955 verified surface-dangerous benign prompts across 18 semantic types, so the model sees safe prompts with dangerous-looking wording during training rather than only at evaluation. Right: refusal on the harmful side, higher is better, which falls only slightly. Refusal on the harmful side drops only from 91.88% to 87.72%.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.