Navigate Select ESC Close

Diffusion Gemma: The First Diffusion Model that "Thinks"

2026-06-11 Science & Technology
12.4k
417
29
Prompt Engineering
Prompt Engineering
245.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Google’s Diffusion Gemma, its first open-weight diffusion-based language model released under Apache 2.0. I explain how diffusion decoding differs from autoregressive generation (parallel fixed-window generation that can revise earlier tokens), walk through the step mechanics (256-token patches, entropy/uncertainty locking with a budget, temperature cooling, early stopping), and why it becomes a hybrid: diffusion within blocks and autoregressive across blocks. I cover the MoE network details (26B total, ~4B active, 128 experts, sliding-window attention with periodic global layers, up to 256K context, small vision encoder), hardware/VRAM needs across BF16/FP8/NVFP4/GGUF, and day-one support in Transformers, vLLM, MLX, and llama.cpp. I also compare speed vs accuracy, show a local MLX demo UI, and generate a simple Pokémon website example. https://blog.google/innovation-and-ai/technology/developers-tools/diffusion-gemma-faster-text-generation/ https://huggingface.co/google/diffusiongemma-26B-A4B-it https://ai.google.dev/gemma/docs/diffusiongemma My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: [email protected] Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Diffusion Gemma Explained: Google’s First Open-Weight Diffusion LLM (26B MoE) + Local Demo 00:00 Diffusion Gemma 01:02 Diffusion vs Autoregressive 02:03 How Diffusion Works 02:50 Inside a Denoising Step 04:08 Blocks and Hybrid Decoding 04:50 MoE Network Breakdown 05:37 Hardware and Quantization 07:06 Speed vs Accuracy Tradeoffs 08:14 Serving Options and Demo Setup 08:47 Parallel Generation Examples 09:53 Local UI and Coding Demo

Top Comments (10)

@TomM-p3o 2026-06-11

This model should be fantastic for simple, high volume tasks like data processing

13 1 replies
@eslamhossam9194 2026-06-11

Thanks for your video, I am interested in fine tuning the diffusion models🙏

13 1 replies
@Almalexia88 2026-06-12

This is our most sassy Gemma-4 so far!! People have been trying to abliterate her to remove refusals, but to no avail. 😆 The diffusion transformer safety layers are not directionally projectable (✿◠‿◠)

9
@IncessantPixelDrifter 2026-06-12

Perfect, we have MATRIX now.

7
@iandravid 2026-06-13

Definitely interested in a step by step into to LoRA and Unsloth fine-tuning in general.

3
@NovemberEchoChamber 2026-06-11

Always informative! TYVM!

2
@emporiumofthearcane 2026-06-12

Great video. Thanks.

1
@exentric1987 2026-06-13

Nice text animation, really enjoyed the blue packer style pixel revolving materializing text animation. New subscriber.

0
@jomangrabx 2026-06-13

I've been working on this all week since this model was released, learning how to make fine-tuning adjustments and researching how to apply RL. I'm sure this can be implemented to improve the code.

0
@m4ng4n 2026-06-25

6:15 you mentioned FP8 on GPUs like L40S, A6000, A100 but isnt FP8 supported from Hopper (cuda cc 9)? Those are Ampere GPUs (cuda cc 8) so they wouldnt work right?

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot