What does the next training paradigm look like?
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
The next paradigm shift (according to Karpathy)
Theo - t3․gg
72.0k views
The data black hole at the center of AI
Dwarkesh Patel
31.1k views
WHCD Shooting: What We Know Now, What’s Next
The Bulwark
187.6k views
Jensen Huang – Will Nvidia’s moat persist?
Dwarkesh Patel
153.3k views
The Department of War is making a huge mistake.
Dwarkesh Patel
35.5k views
Dario Amodei — “We are near the end of the exponential”
Dwarkesh Patel
41.0k views
The US v Iran: What does war look like? | The Security Brief
BBC News
287.0k views
THIS Is What a Strong Democrat Looks Like | The Next Level
The Bulwark
217.0k views
If I Started Running Today, This Is What I’d Do & Buy!
Ben Parkes
253.4k views
What are we scaling?
Dwarkesh Patel
22.3k views
Top Comments (10)
Dwarkesh RL’ed me into checking YouTube two minutes after he posts a video and hitting the like button.
Love these audio essays. Improves my understanding massively
I think the biggest issue is probably organizational IP, etc. Personal users might be fine training and contributing the bigger model, but companies want to protect their IP. It might end up being "company specific AI", where for a given company, an AI has access to, and can learn. What this looks like is probably still a frozen base model (maybe trained through personal users that share their sessions), and maybe LoRA-like things for personal users in the interim while it hasn't completely learned. And then take the personal LoRAs and integrate them into the big ai model during /dream, and then for the company, into a bigger LoRA for that company specifically.
Even organizations don’t know how they work and how anything gets done
How dare the algorithm hide this video from my for 34 seconds.
Perhaps an interview with Yann Lecun? He has a different view on the direction of AI with recently introduced “World Models” based on Joint Embedding Predictive Architecture (JEPA) I believe he and his team introduced CNNs in 1990s
One aspect of this "computationally-bound learnability" is a direct description of Finzi et, al's "Epiplexity" paper, highly worth getting them on the pod to discuss this.
That's why your ai describes the problem, then grinds, and finally summarizes. The summary is the training chow for the next gen of models. "What I should have done" is the same as "what the next generation should do in the first place". Write memoirs and build a search index from them.
I think the key to sample efficiency is somehow encouraging and/or encoding “emotion” in the models. I’m “sample efficient” because I forget the vast majority of my experiences… I just remember the highs and lows. Which for me are geared around me surviving and reproducing. I guess what’s interesting is that somehow my brain solves the credit assignment problem for me… actions that I take that harm my chances at prolifically spreading my genes feel “bad”, even if they’re extremely far removed from the source signal. Like of course almost dying feels bad, and you’ll remember that forever, but that seems relatively easy to implement… less obvious is stuff like feeling bad if I don’t comb my hair, because then others will perceive me as a slob, and girls won’t find that attractive, and, etc, etc… marginally diminished chance of having babies. We would need something that understands “this action anomalously diminished our chances of success” and “this action anomalously increased our chances”… and with those signals, chew on the highs and lows to figure out “why did this action change my chances of success?” … I suppose though that by definition requires something “wiser” than the thing being trained that can guide it. In a sense, our instincts are “wiser” than we are because baked into them are millions of years of experience toward accomplishing a simple task. That’s probably a lot harder for AI when we’re constantly throwing new tasks at it… especially tasks we’ve never seen successfully completed
6:00 "What is the RL environment for building an AI as good at politics as Lyndon Johnson or as good at building a space launch business as Elon Musk?" Bravo Dwarkesh. Brilliantly framed. You are asking the real questions here. Very subtle :)
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
Dwarkesh RL’ed me into checking YouTube two minutes after he posts a video and hitting the like button.
Love these audio essays. Improves my understanding massively
I think the biggest issue is probably organizational IP, etc. Personal users might be fine training and contributing the bigger model, but companies want to protect their IP. It might end up being "company specific AI", where for a given company, an AI has access to, and can learn. What this looks like is probably still a frozen base model (maybe trained through personal users that share their sessions), and maybe LoRA-like things for personal users in the interim while it hasn't completely learned. And then take the personal LoRAs and integrate them into the big ai model during /dream, and then for the company, into a bigger LoRA for that company specifically.
Even organizations don’t know how they work and how anything gets done
How dare the algorithm hide this video from my for 34 seconds.
Perhaps an interview with Yann Lecun? He has a different view on the direction of AI with recently introduced “World Models” based on Joint Embedding Predictive Architecture (JEPA) I believe he and his team introduced CNNs in 1990s
One aspect of this "computationally-bound learnability" is a direct description of Finzi et, al's "Epiplexity" paper, highly worth getting them on the pod to discuss this.
That's why your ai describes the problem, then grinds, and finally summarizes. The summary is the training chow for the next gen of models. "What I should have done" is the same as "what the next generation should do in the first place". Write memoirs and build a search index from them.
I think the key to sample efficiency is somehow encouraging and/or encoding “emotion” in the models. I’m “sample efficient” because I forget the vast majority of my experiences… I just remember the highs and lows. Which for me are geared around me surviving and reproducing. I guess what’s interesting is that somehow my brain solves the credit assignment problem for me… actions that I take that harm my chances at prolifically spreading my genes feel “bad”, even if they’re extremely far removed from the source signal. Like of course almost dying feels bad, and you’ll remember that forever, but that seems relatively easy to implement… less obvious is stuff like feeling bad if I don’t comb my hair, because then others will perceive me as a slob, and girls won’t find that attractive, and, etc, etc… marginally diminished chance of having babies. We would need something that understands “this action anomalously diminished our chances of success” and “this action anomalously increased our chances”… and with those signals, chew on the highs and lows to figure out “why did this action change my chances of success?” … I suppose though that by definition requires something “wiser” than the thing being trained that can guide it. In a sense, our instincts are “wiser” than we are because baked into them are millions of years of experience toward accomplishing a simple task. That’s probably a lot harder for AI when we’re constantly throwing new tasks at it… especially tasks we’ve never seen successfully completed
6:00 "What is the RL environment for building an AI as good at politics as Lyndon Johnson or as good at building a space launch business as Elon Musk?" Bravo Dwarkesh. Brilliantly framed. You are asking the real questions here. Very subtle :)