AI article

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

No preview is available. Read the original article for the full story.

huggingface | Sep 8, 2026 | Antonio Tiene, Alejo Lopez Avila, Iker García-Ferrero

Automated excerpt

So the real problem is not only raising refusal on harmful prompts, but shaping the behaviour near the boundary itself. Safety tuning tends to produce false refusals on benign prompts that look superficially dangerous. To compensate, we build in-distribution benign data, including 11,955 verified surface-dangerous benign prompts across 18 semantic types, so the model sees safe prompts with dangerous-looking wording during training rather than only at evaluation. Right: refusal on the harmful side, higher is better, which falls only slightly. Refusal on the harmful side drops only from 91.88% to 87.72%.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news