Tech article
GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt
No preview is available. Read the original article for the full story.
Hacker News | Sep 15, 2026 | vital101
Automated excerpt
Despite extensive work on safety post-training, it has been shown that models can be readily unaligned through post-deployment fine-tuning. Moreover, GRP-Oblit generalizes beyond language models and can also unalign diffusion-based image generation systems. We evaluate GRP-Oblit on six utility benchmarks and five safety benchmarks across fifteen 7-20B parameter models, spanning instruct and reasoning models, as well as dense and MoE architectures.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.