AI article

RoPE vs sinusoidal positional encoding, measured: 55 logits of drift against 0.0005

Community description: Take two token embeddings. Keep them exactly five positions apart and slide the pair down...

Dev.to | Oct 1, 2026 | Mira Ceti

Automated excerpt

Under RoPE the same pair scores -0. 6102 every single time, to within no algebra. Every number on this page was before publishing, not copied from a paper. x_q, x_k = torch. randn(D), torch. randn(D) # two token embeddings Wq, Wk = torch. randn(D, D) / D**0. 5, torch. randn(D, D) / D**0. 5 return ((x_q + sinu. pe[m]) @ Wq) @ ((x_k + sinu. pe[n]) @ Wk) at = lambda x, W, p: rope((x @ W). view(1, 1, 1, D), seq_positions=torch. tensor([[p]])) return (at(x_q, Wq, m) * at(x_k, Wk, n)).

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news