Navigate Select ESC Close

The NEW Geometry Behind AI Reasoning (Princeton, Berkeley)

2026-07-03 Science & Technology
1.9k
95
16
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Modern Transformers don't simply fail because they forget: they often fail because their internal representations drift away from the geometry that later computations expect. In this video, we'll dissect the mechanistic interpretability behind "DiscoLoop", follow the causal experiments that reveal the hidden representation mismatch inside recurrent Transformers, and see why a simple embedding-alignment principle dramatically improves multi-hop reasoning and out-of-distribution generalization. If you're interested in latent reasoning, representation geometry, recurrent architectures, and the future of AI cognition, this paper offers one of the most elegant architectural insights of the year. all rights w/ authors: DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning" Hengyu Fu1∗ Tianyu Guo1∗ Zixuan Wang1,2∗ Hanlin Zhu1∗ Jason D. Lee1 Jiantao Jiao1 Stuart Russell1 Song Mei1 from 1 University of California, Berkeley. 2 Princeton University (!) #airesearch #nextgenai #newtechnology #artificialintelligence

Top Comments (8)

@TheAdeybob 2026-07-03

Been wondering about getting past the noise in order to give better hopping outcomes. My solution was similar to the Discoloop team's ideas. All in all tho, this DiscoLoop idea makes me wonder if their process can also be externalised. I mean, apart from some tinkering, there's nowt stopping an LLM or wrapper handing its current reasoning signal/context to a deterministic layer, which canonicalises it, audits it, vectorises useful chunks - and when it's finished mucking about, then reinjecting the cleaned signal before the next hop or final answer. Could save a whole lot of problems, and tokens...and much of the compute is handed off to the user's local resources. Anyhows, with open-weight models, hooks could make this direct. But even black-box LLMs could use a weaker version through structured context handoff and reinjection. DiscoLoop may point at some cool proofs - not only regards internal recurrence, but also regards external audit-reinjection circuits that're usable around almost any coherent LLM. Depends how you squint at it.

3 3 replies
@seanoconnor1984 2026-07-03

You are just doing everything super difficult. The recipe is binary context and an analog feature vector flowing through context dependent linear mappings. So for ReLU networks the binary context is the local x>=0? decisions and then you just have the feature vector flowing through the selected mapping at each layer. It need not be that way though and it can be a lot simpler. There is a thing now, Atlas LSH neural networks or even more general Atlas neural networks.

2 5 replies
@mitchellwilley7208 2026-07-03

love ur videos, this is actually something i have been working on myself.

1 2 replies
@timmygilbert4102 2026-07-03

I actually don't remember if i commented this on the jepas video or just jolt the note, re pasting an extract: "The idea is that word token act as a regularisation that snap back to neutral center of mass of a particular meaning, and the interpolation from attention conserve both smooth signal and relative positioning to centers of meaning through voronoi like domain, which jepa lack." They brought that observation to the internal, which is something i wanted to try but in a different way -> (internal concept token learning) It would solve mire than two hop reasoning, but help with tunneling issues too. And lay the fundation for a type of neuro symbolique reasoning approach through discrete concept landmark. Glad my intuition are still right. But that mean that quantisation at training pass might help given right enough resolution. I just don't have explored training dynamic relative to first principles to figure out how to handle complex geometric learning, yet, ly mental model is heavily forward like, which is my weakness.

0
@TheAdeybob 2026-07-03

Excellent vid. Looking forward to the next one!

0
@peterbabu936 2026-07-03

who remembers DYTOPO?

0
@bjmay67 2026-07-03

Love this approach! Seems to work for their tests... will take a look at the paper. Thanks!

0
@tantzer6113 2026-07-04

Thank you! I love videos on architectures that employ recurrence.

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot