Loading…

Instead of training a new model from scratch, Google DeepMind retrofitted Gemma 4 into a diffusion model using less than 10 percent of the original training budget. DiffusionGemma generates 256 tokens in parallel instead of one at a time, hitting about 1,500 tokens per second.…
To respect copyright, we link to the source rather than republishing the full text. Read the complete article on The Decoder.