Tech article
Training Text-to-Image Models Without a VAE
No preview is available. Read the original article for the full story.
Hacker News | Oct 6, 2026 | schopra909
Automated excerpt
Only four Linum v2 image checkpoints survive, so its curve starts at 117M samples. All three gated residual experiments have the same 11,776-dimensional residual stream (4 × 2,944). Block 10 predicts a 128×128 image against x₀ downsampled 4×, block 16 a 256×256 image against x₀ downsampled 2×, and the final head the full 512×512 image.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.