Tech article
Training Text-to-Image Models 3.6× Faster
No preview is available. Read the original article for the full story.
Hacker News | Sep 16, 2026 | schopra909
Automated excerpt
Almost all generative image and video models are Latent Diffusion Models (LDMs). JiT (wide) · 512×5122. 0B active pixel-space DiT256 pixel tokensJiT-DDT 64/32 (baseline) · 512×5122. 2B active pixel-space DiT320 pixel tokens = 64 encoder + 256 decoderBoth models are trained on the same 100M samples. They then train two independent models, the flow matching generative model and a decoder from DINO space back to pixel space.
Selected automatically from source text; not independently written or fact-checked. Read the original for full context.