Navigate Select ESC Close

The End of External World Models: Folded into LLMs #ai

2026-06-30 Science & Technology
1.7k
77
10
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

"All rights w/ authors: Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning" Xuan Zhang1,2,3, Zhijian Zhou1,2,3, Lingfeng Qiao3, Yulei Qin3, Ke Li3, Xing Sun3, Xiaoyu Tan3†, Chao Qu1†, Yuan Qi1† from 1 Fudan University, 2 Shanghai Innovation Institute, 3 Tencent Youtu Lab #airesearch #aiexplained #scienceexplained #scienceexperiment #futuretech #aiagents

Top Comments (6)

@tom-et-jerry 2026-06-30

If the world model is linguistic, spontaneous reflexes would not be possible in the case of an embodied robot. I fail to see why we cannot establish a latent space for a world model.

4 2 replies
@stormbreaker420 2026-06-30

the world model saga has begun

2 2 replies
@timmygilbert4102 2026-06-30

Mtp with extra steps

2 2 replies
@Dale-n4n 2026-07-01

Maybe i misunderstood something, but few thoughts: - Let's say we humans learn to play with (and within) the physical reality and we store our experiences (positive/negative) as "trajectories" in our memory in a neurochemical pathways, but when we try to think about our next step we retrieve our memories into "active memory" (or imagination space) as words and visuals (eg. short reels). - So, our learned trajectories shape our percieved "description of reality" from a set of stories. - We believe that we've recorded "facts" about reality, but then a new story happens (in our or other's experience) which slightly or radically changes our "understanding" (ie. lingo-visual description) of previous memories, eg. "the sweet and tasty CAN be deadly", or "the Sun does NOT revolve around Earth". - When we play a game or a musical instrument we accumulate many "tried trajectories" also as a "muscle memory" (whatever that means, but which is *not just* lingo-visual). So, seems to me that this WM-AMT approach could be improved in two ways: - add a self-enhancement mechanism: a model or harness should re-rank it's weights when new facts or trajectories become relevant to the old facts/trajectories learned at the mid-training stage. - we can use other techniques/WM like unsupervised training (eg. games/simulations) combined with linguistic descriptions of such experiments (about each environment, actions, outcomes) to synthesize more "plausible trajectories" for WM-AMT for niche applications/problem spaces. Ofc, this would require more compute resources.

2 1 replies
@bjmay67 2026-06-30

33:43 I personally don't think that LLMs "have" an implicit WM, at least not a sufficiently rich one. At best they have a sparse Linguistic (highly compressed/abstracted) Model of certain probabilistically connected text trajectories (rollots) - which can be "enhanced" a bit with training to act as if it "knows" certain world state transitions. ("Summaries" as you say.) OTOH, non-language transformers trained on highly specialized tokens / vectors / tensors other than language (actual parameter states) could learn some of the state-space manifold and its transitions, from a probabilistic POV, like a Bayesian S-matrix. But then why not use Bayesian or other networks and skip LLM wishfull thinking? It seems many people "project" LLMs should be cognitively able do anything and everything, while the evidence suggests they are rather fragile outside of their strengths... hence 100s of training methodologies, harnesses, validity checkers, etc.

1
@jtjames79 2026-06-30

Every word is a name for something. It's all name magic. 🤔 Any sufficiently advanced magic is indistinguishable from technology. 🤷

1

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot