AI article
Training a model to identify AI-generated web content from structure alone
No preview is available. Read the original article for the full story.
Hacker News | Sep 22, 2026 | jochenmadler
Automated excerpt
Abstract:Word-level detectors identify unedited AI-generated text almost perfectly, but the literature documents their brittleness under rewording, and a word-level score neither characterizes a text nor identifies which AI model wrote it. We replicate StoryScope (Russell et al. , 2026), which showed such patterns for AI-generated fiction, on commercial content: 2,250 pre-ChatGPT human blog posts from 268 company domains against 11,250 AI mirrors from five frontier models. The signal characterizes and attributes: AI posts share a tidy, self-announcing shape, 79. 3% are attributed to the correct source against a 16. 7% chance rate, and human posts occupy rare structural configurations.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.