AI article

25 LLM architecture blocks, side by side, in runnable PyTorch

Community description: GPT-2 to Kimi Linear is seven years of architecture research, and almost all of it fits in about...

Dev.to | Sep 24, 2026 | Mira Ceti

Automated excerpt

Llama 2, Llama 3, Llama 4, Qwen 2. 5, Phi-3, Phi-4, Mistral Small 3. 1, Nanbeige 4. 1 Gemma 2, Qwen 3, Qwen3-Next, Qwen3. 5, OLMo 3, DeepSeek-V3, MiniMax-M2, MiniMax-M2. 5, Mistral Large 3, Kimi K2, Kimi Linear, Ling 2. 5, Sarvam 30B return torch. tanh(logits / self. softcap) * self. softcap Embedding scaling by sqrt(d) and tanh soft-capping of logits (30. 0 final, 50. 0 on attention logits). Attention: four families, not one GPT-2, OPT, OLMo; Llama 2 7B/13B, Phi-3. 5 Mini Llama 3, Qwen 2. 5/3, Gemma 2, Phi-4, Mistral Small 3. 1, OLMo 3, MiniMax-M2/2.

Selected automatically from source text; not independently written or fact-checked. Read the original for full context.

Read the original article

More AI news