Navigate Select ESC Close

Diffusion Is Coming for Text. Here's NVIDIA's New Model.

2026-07-06 Science & Technology
3.3k
104
13
Prompt Engineering
Prompt Engineering
245.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

NVIDIA’s Two-Tower Diffusion Language Model (Nemotron): Faster Text Generation with Frozen Context In this video, I break down how diffusion is moving into text generation and why it can avoid the compute and memory limits of autoregressive next-token prediction by generating in parallel. I explain NVIDIA’s Nemotron “Two Tower” diffusion language model: two cloned 52-layer towers (Mamba-2 + self-attention + MoE), where one tower is frozen as a left-to-right context model and the other is retrained as a denoiser that fills masked “noise” in 16-token blocks. I cover the layer-by-layer cross-attention “sky bridges,” the diffusion-style timer add-on, quality retention (about 98.7% vs the original), benchmark tradeoffs (math/code drops), ablation results showing freezing is key, and brittleness when changing block size (16 to 64 collapses generation). @NVIDIADeveloper LINKS: https://huggingface.co/nvidia/Nemotron-Labs-TwoTower-30B-A3B-Base-BF16 My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: [email protected] Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 00:00 Diffusion Nemotron 01:03 Two Tower Big Idea 02:25 Architecture and Benchmarks 03:58 Why Two Towers Work 06:14 Blockwise Diffusion Decoding 09:31 Limits and What’s Next

Top Comments (10)

@iandravid 2026-07-06

didn’t diffusion-gemma do the same thing with the 26b MoE LLM as base, but completely changing the autoregressive output step to diffusion-based block generation? just confirming if they are the same approach?

11 3 replies
@nguyenanhnguyen7658 2026-07-06

NVDIA has NEVER EVER made a reasonable model.

1 2 replies
@arx6.9 2026-07-06

Wait isn't this just like speculative decoding using smaller drafter models ?

1 2 replies
@AkashSwamyBazinga 2026-07-06

seems like this is not going to scale well, is 30B splits into 2 towers or 30Bx2? how does it stack against a let's say 60B model?

1 1 replies
@mevech 2026-07-06

Would be nice to have diffusion with myp right?

0 1 replies
@simone_rizzo98 2026-07-07

hi how do you make this animations with blackboard?

0 1 replies
@HacknSlashPro 2026-07-07

already saw diffusion based gemma, it was producing exponentially more inaccurate output compared to transformer base version

0 1 replies
@abhirj87 2026-07-09

inception labs mercury 2 is also on similar lines. Token per sec is extreemly fast. The model is not bad and costs 1/5th of GPT 5.2

0
@NVIDIADeveloper 2026-07-08

thanks for sharing! 🙌

0
@bobharris5093 2026-07-06

i failed to be excited

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot