Navigate Select ESC Close

What does the next training paradigm look like?

2026-06-26 Science & Technology
49.5k
1.8k
191
Dwarkesh Patel
Dwarkesh Patel
1.4m subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Thanks to Mercury for sponsoring this essay. Mercury has automated basically my entire bill pay process for my business. I just give contractors a dedicated email address, and when they send an invoice, Mercury automatically creates a draft payment for me to review. I no longer have to hunt through my inbox for invoices or deal with messy spreadsheets to track my bills. Mercury handles it all. Learn more at https://mercury.com 𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 Read the essay here: https://www.dwarkesh.com/p/the-next-paradigm Sasha Rush lecture: https://youtu.be/wxOZWD6wYVY TIMESTAMPS 00:00:00 – The big research bet the labs are making 00:02:12 – Grindability is just as important as verifiability 00:06:10 – Will RLVR alone generalize? 00:08:41 – Getting the learning back to the weights 00:15:22 – Dreaming 00:17:23 – What 2027 looks like

Top Comments (10)

@midsingularity 2026-06-26

Dwarkesh RL’ed me into checking YouTube two minutes after he posts a video and hitting the like button.

84
@brianandrewstuart 2026-06-26

Love these audio essays. Improves my understanding massively

67 2 replies
@howdocowsfly 2026-06-26

I think the biggest issue is probably organizational IP, etc. Personal users might be fine training and contributing the bigger model, but companies want to protect their IP. It might end up being "company specific AI", where for a given company, an AI has access to, and can learn. What this looks like is probably still a frozen base model (maybe trained through personal users that share their sessions), and maybe LoRA-like things for personal users in the interim while it hasn't completely learned. And then take the personal LoRAs and integrate them into the big ai model during /dream, and then for the company, into a bigger LoRA for that company specifically.

56 8 replies
@scarface548 2026-06-26

Even organizations don’t know how they work and how anything gets done

48 1 replies
@JordanLynn 2026-06-26

How dare the algorithm hide this video from my for 34 seconds.

37 1 replies
@chamathamara 2026-06-26

Perhaps an interview with Yann Lecun? He has a different view on the direction of AI with recently introduced “World Models” based on Joint Embedding Predictive Architecture (JEPA) I believe he and his team introduced CNNs in 1990s

21 1 replies
@newplace2frown 2026-06-26

One aspect of this "computationally-bound learnability" is a direct description of Finzi et, al's "Epiplexity" paper, highly worth getting them on the pod to discuss this.

17 1 replies
@easydoesitismist 2026-06-26

That's why your ai describes the problem, then grinds, and finally summarizes. The summary is the training chow for the next gen of models. "What I should have done" is the same as "what the next generation should do in the first place". Write memoirs and build a search index from them.

6
@ryanfranz6715 2026-06-26

I think the key to sample efficiency is somehow encouraging and/or encoding “emotion” in the models. I’m “sample efficient” because I forget the vast majority of my experiences… I just remember the highs and lows. Which for me are geared around me surviving and reproducing. I guess what’s interesting is that somehow my brain solves the credit assignment problem for me… actions that I take that harm my chances at prolifically spreading my genes feel “bad”, even if they’re extremely far removed from the source signal. Like of course almost dying feels bad, and you’ll remember that forever, but that seems relatively easy to implement… less obvious is stuff like feeling bad if I don’t comb my hair, because then others will perceive me as a slob, and girls won’t find that attractive, and, etc, etc… marginally diminished chance of having babies. We would need something that understands “this action anomalously diminished our chances of success” and “this action anomalously increased our chances”… and with those signals, chew on the highs and lows to figure out “why did this action change my chances of success?” … I suppose though that by definition requires something “wiser” than the thing being trained that can guide it. In a sense, our instincts are “wiser” than we are because baked into them are millions of years of experience toward accomplishing a simple task. That’s probably a lot harder for AI when we’re constantly throwing new tasks at it… especially tasks we’ve never seen successfully completed

3
@TehNarrator 2026-06-27

6:00 "What is the RL environment for building an AI as good at politics as Lyndon Johnson or as good at building a space launch business as Elon Musk?" Bravo Dwarkesh. Brilliantly framed. You are asking the real questions here. Very subtle :)

2

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot