The End of External World Models: Folded into LLMs #ai
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
The Most Exciting Discoveries | The Curse of Oak Island
HISTORY
60.7k views
Major Discovery on the Origin of Life Found Inside a Korean Crater
Anton Petrov
47.2k views
This model is kind of a disaster.
Theo - t3․gg
166.2k views
Possible Discovery of First Ever Stars in the Universe
Anton Petrov
73.9k views
New Start Treaty must be extended or else world in danger
The Duran
20.7k views
The Highest Protein, Most Anti-Inflammatory Nut in the World has Been Discovered
Thomas DeLauer
31.2k views
Supreme Court In a WORLD OF FEAR after Disaster They CREATED
Legal AF
126.3k views
Trump Makes Enemies and Buys Friends All Over the World | The Daily Show
The Daily Show
1.2m views
Modern Marvels: Engineering Disasters That Changed the World Forever (S9, E19) | Full Episode
HISTORY
255.8k views
You’re Watching the End of the World in Real Time - Eric Weinstein
The Diary Of A CEO
5.1m views
Top Comments (6)
If the world model is linguistic, spontaneous reflexes would not be possible in the case of an embodied robot. I fail to see why we cannot establish a latent space for a world model.
the world model saga has begun
Mtp with extra steps
Maybe i misunderstood something, but few thoughts: - Let's say we humans learn to play with (and within) the physical reality and we store our experiences (positive/negative) as "trajectories" in our memory in a neurochemical pathways, but when we try to think about our next step we retrieve our memories into "active memory" (or imagination space) as words and visuals (eg. short reels). - So, our learned trajectories shape our percieved "description of reality" from a set of stories. - We believe that we've recorded "facts" about reality, but then a new story happens (in our or other's experience) which slightly or radically changes our "understanding" (ie. lingo-visual description) of previous memories, eg. "the sweet and tasty CAN be deadly", or "the Sun does NOT revolve around Earth". - When we play a game or a musical instrument we accumulate many "tried trajectories" also as a "muscle memory" (whatever that means, but which is *not just* lingo-visual). So, seems to me that this WM-AMT approach could be improved in two ways: - add a self-enhancement mechanism: a model or harness should re-rank it's weights when new facts or trajectories become relevant to the old facts/trajectories learned at the mid-training stage. - we can use other techniques/WM like unsupervised training (eg. games/simulations) combined with linguistic descriptions of such experiments (about each environment, actions, outcomes) to synthesize more "plausible trajectories" for WM-AMT for niche applications/problem spaces. Ofc, this would require more compute resources.
33:43 I personally don't think that LLMs "have" an implicit WM, at least not a sufficiently rich one. At best they have a sparse Linguistic (highly compressed/abstracted) Model of certain probabilistically connected text trajectories (rollots) - which can be "enhanced" a bit with training to act as if it "knows" certain world state transitions. ("Summaries" as you say.) OTOH, non-language transformers trained on highly specialized tokens / vectors / tensors other than language (actual parameter states) could learn some of the state-space manifold and its transitions, from a probabilistic POV, like a Bayesian S-matrix. But then why not use Bayesian or other networks and skip LLM wishfull thinking? It seems many people "project" LLMs should be cognitively able do anything and everything, while the evidence suggests they are rather fragile outside of their strengths... hence 100s of training methodologies, harnesses, validity checkers, etc.
Every word is a name for something. It's all name magic. 🤔 Any sufficiently advanced magic is indistinguishable from technology. 🤷
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (6)
If the world model is linguistic, spontaneous reflexes would not be possible in the case of an embodied robot. I fail to see why we cannot establish a latent space for a world model.
the world model saga has begun
Mtp with extra steps
Maybe i misunderstood something, but few thoughts: - Let's say we humans learn to play with (and within) the physical reality and we store our experiences (positive/negative) as "trajectories" in our memory in a neurochemical pathways, but when we try to think about our next step we retrieve our memories into "active memory" (or imagination space) as words and visuals (eg. short reels). - So, our learned trajectories shape our percieved "description of reality" from a set of stories. - We believe that we've recorded "facts" about reality, but then a new story happens (in our or other's experience) which slightly or radically changes our "understanding" (ie. lingo-visual description) of previous memories, eg. "the sweet and tasty CAN be deadly", or "the Sun does NOT revolve around Earth". - When we play a game or a musical instrument we accumulate many "tried trajectories" also as a "muscle memory" (whatever that means, but which is *not just* lingo-visual). So, seems to me that this WM-AMT approach could be improved in two ways: - add a self-enhancement mechanism: a model or harness should re-rank it's weights when new facts or trajectories become relevant to the old facts/trajectories learned at the mid-training stage. - we can use other techniques/WM like unsupervised training (eg. games/simulations) combined with linguistic descriptions of such experiments (about each environment, actions, outcomes) to synthesize more "plausible trajectories" for WM-AMT for niche applications/problem spaces. Ofc, this would require more compute resources.
33:43 I personally don't think that LLMs "have" an implicit WM, at least not a sufficiently rich one. At best they have a sparse Linguistic (highly compressed/abstracted) Model of certain probabilistically connected text trajectories (rollots) - which can be "enhanced" a bit with training to act as if it "knows" certain world state transitions. ("Summaries" as you say.) OTOH, non-language transformers trained on highly specialized tokens / vectors / tensors other than language (actual parameter states) could learn some of the state-space manifold and its transitions, from a probabilistic POV, like a Bayesian S-matrix. But then why not use Bayesian or other networks and skip LLM wishfull thinking? It seems many people "project" LLMs should be cognitively able do anything and everything, while the evidence suggests they are rather fragile outside of their strengths... hence 100s of training methodologies, harnesses, validity checkers, etc.
Every word is a name for something. It's all name magic. 🤔 Any sufficiently advanced magic is indistinguishable from technology. 🤷