Navigate Select ESC Close

The data black hole at the center of AI

2026-06-19 Science & Technology
31.1k
1.6k
180
Dwarkesh Patel
Dwarkesh Patel
1.4m subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

It is easy to forget how much data these models are trained on, and how much more it is than what we humans see in our lifetimes. We see these AIs as a galaxy glittering with capabilities, but at their center, invisible to the naked eye, holding all the constellations together, is an unimaginably massive black hole of data. Thanks to Mercury for sponsoring this essay. Mercury is my banking platform, and they just released a new AI feature called Command. Since I already use Mercury to run basically my entire business, Command has access to all the info it needs to get real work done. I can ask it to send invoices, or categorize expenses, or even transfer money… and Command just handles it. Learn more at https://mercury.com/command Read the transcript here: https://www.dwarkesh.com/p/the-sample-efficiency-black-hole-2 TIMESTAMPS 00:00:00 – What is really driving AI progress? 00:03:11 – Comparing human vs AI sample efficiency 00:08:46 – Does sample efficiency matter?

Top Comments (10)

@davidbutler9323 2026-06-19

My baby had seen three sinks before she knew what the completely different looking sink at the library was.

105 5 replies
@ericadar 2026-06-19

The objection to Karpathy's argument regarding evolution pre-training may be incomplete. The parameter encoding may not be restricted to the genome. There may be significant encoding in the geometry of the gamete, the distribution of its organelles, the enzymes, and other data-encoding molecules. Without the initial machinery, the genome is not sufficient to instantiate biological intelligence.

88 9 replies
@sohamgawand 2026-06-19

dwarkesh, its about time for this year's Trenton and Sholto podcast

47 3 replies
@anusmith 2026-06-19

Maybe the real ASI was the friends we made along the way...

36 1 replies
@slash196 2026-06-19

4:47 You should remember the interview you JUST did with David Reich where he pointed out the epigentic coding the tells our genome the context in which any given gene should be expressed which amplifies the computational complexity of the genome MASSIVELY.

31
@ECTCa 2026-06-19

It seems like you will be getting Yann LeCun on your podcast. Your thesis here aligns with his vision on the future of AI

22 3 replies
@rodrimora 2026-06-19

"We have models that are 5T parameters" >What did Dwarkesh saw?

21 3 replies
@basils.254 2026-06-19

Dwarkesh finally interviews himself and it's his best video yet

20
@PhotoninDark 2026-06-19

The next big leap for AI is to reach at least 0.000001 of the learning efficiency of the Human Brain.

14
@wren1728 2026-06-19

I think we should probably be very concerned about this sample efficiency gap from a safety perspective. Supposing that the labs are correct that their automated R&D could realise algorithms with human level sample efficiency, then we might very suddenly move to a world where the relative quantity of available data jumps by orders of magnitude without much warning. What would it be like if a human could spend thousands of years learning? We have no idea.

6 1 replies

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot