Tech article
A study of sequence weighting at scale
No preview is available. Read the original article for the full story.
Hacker News | Sep 21, 2026 | pranitha_m
Automated excerpt
Taken together, our results are consistent with a non-monotonic rise-then-fall in the effective sequence weight exponent: small-scale models fit small effective sequence weight exponents, learning patterns across the entire dataset independent of data weight; medium-scale models fit larger effective sequence weight exponents, learning patterns in data proportional to their data weight. Large-scale models once again fit small effective sequence weight exponents, learning all patterns present in the data regardless of weight. Epoching shifts the effective sequence weight peak towards smaller models.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.