Diffusion Is Coming for Text. Here's NVIDIA's New Model.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Software engineering is dead now
Theo - t3․gg
165.4k views
Sonnet 4.5 Is Here—And It’s a Beast at Coding
Prompt Engineering
52.0k views
Python is Changing – Here’s What’s Coming
Tech With Tim
165.5k views
Context Engineering is All You NEED!
Prompt Engineering
38.7k views
Context Engineering for Agents
LangChain
189.5k views
The Only Embedding Model You Need for RAG
Prompt Engineering
35.2k views
Context Engineering is the future of AI Agents - here’s why
David Ondrej
44.4k views
Gemini CLI — Google’s Free Open-Source Coding Agent
Prompt Engineering
56.6k views
The New World Economy: How to Prepare for What’s Coming
Minority Mindset
233.4k views
The Secret to Perfect Prompts (Without Prompt Engineering)
Futurepedia
55.4k views
Top Comments (10)
didn’t diffusion-gemma do the same thing with the 26b MoE LLM as base, but completely changing the autoregressive output step to diffusion-based block generation? just confirming if they are the same approach?
NVDIA has NEVER EVER made a reasonable model.
Wait isn't this just like speculative decoding using smaller drafter models ?
seems like this is not going to scale well, is 30B splits into 2 towers or 30Bx2? how does it stack against a let's say 60B model?
Would be nice to have diffusion with myp right?
hi how do you make this animations with blackboard?
already saw diffusion based gemma, it was producing exponentially more inaccurate output compared to transformer base version
inception labs mercury 2 is also on similar lines. Token per sec is extreemly fast. The model is not bad and costs 1/5th of GPT 5.2
thanks for sharing! 🙌
i failed to be excited
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
didn’t diffusion-gemma do the same thing with the 26b MoE LLM as base, but completely changing the autoregressive output step to diffusion-based block generation? just confirming if they are the same approach?
NVDIA has NEVER EVER made a reasonable model.
Wait isn't this just like speculative decoding using smaller drafter models ?
seems like this is not going to scale well, is 30B splits into 2 towers or 30Bx2? how does it stack against a let's say 60B model?
Would be nice to have diffusion with myp right?
hi how do you make this animations with blackboard?
already saw diffusion based gemma, it was producing exponentially more inaccurate output compared to transformer base version
inception labs mercury 2 is also on similar lines. Token per sec is extreemly fast. The model is not bad and costs 1/5th of GPT 5.2
thanks for sharing! 🙌
i failed to be excited