Open 32B w/ AutoMemory beats Opus: HOW? (Stanford)
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Stanford Neuroscientist: Can’t Remember Your Dreams? Your Brain May Be Warning You!
The Diary Of A CEO
27.5k views
2025’s Biggest Alien Discoveries (Part 2) | Ancient Aliens
HISTORY
106.8k views
Trump's White House DISASTER gets Blown WIDE OPEN!
Adam Mockler
57.4k views
Running Injury? Recover Quicker and Stay Fit... Here's How!
Ben Parkes
38.4k views
The Truth about AI is Devastating: Proof by MIT, Harvard
Discover AI
65.2k views
AI is Burning - 5 New Papers
Discover AI
34.3k views
How I'd Learn ML/AI FAST If I Had to Start Over
Tech With Tim
289.9k views
OpenAI's "AI SYSTEMS" and New Scientific Discoveries
Wes Roth
47.9k views
OpenAI's Autonomous AI Research Benchmark
Wes Roth
48.3k views
Discovered: Top Secret Alien Bases | Ancient Aliens
HISTORY
854.0k views
Top Comments (10)
Interesting, useful, but we need methods that are not dependent on frontier models
It’s a bit uncanny how often your videos feel like they are crawling my recent agent logs to explain and expand the exact thing I’m working on at the time. 😂
So rather than just using Claude, I can now use Claude and Qwen and distill Claude knowledge into a LoRA for Qwen, and then I basically get near the same intelligence as Claude, provided that I keep interacting with the model about the same thing, and I only get this experience after some usage and training. Or ... I could have just used Claude. Sure, Claude is money, and Qwen is free, but even with these training loops, I still have to pay for Claude. If I have a Toyota car, and I want it to run better, it's like I'm being told to buy a Ferrari, take pieces out of the engine of the Ferrari, and put them into the Toyota, and then the Toyota will nearly run as good as the Ferrari. But if I bought a Ferrari, why wouldn't I just drive the Ferrari and ditch the Toyota all together? All this is is perpetual distillation, in which I pay for Claude to distill to Qwen to almost function at a Claude level. Interesting that it works, but I have better thing to waste my time with that provide a higher value of return. Not to mention, what about all the time I use Qwen pre-training, the wait time during traing, the post-trained usage where I'm confident in Qwen's results, yet might be engaging it on a topic it hadn't yet been trained and get a bad answer. None of this is useful.
What are the best Hermes agent memory addons? Is obsidian enough?
Can you check out the StateLM and CaveAgent papers?
It's LLMs all the way down.
More good stuff here. 👏
One of your future videos will be about abstracting and three piece or more structures broken into math.... 🙂
Can't wait to see when 1B parameter model with well trained memory skill will outperform all of the beefy Terra-scale models, that require ton of memory and compute to run. This is really way forward - teach a model how to condense and distill the knowledge from the noise, instead of tuning billions of parameters by treating noise as 'useful data to learn' during training.
More epistemic foraging!
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
Interesting, useful, but we need methods that are not dependent on frontier models
It’s a bit uncanny how often your videos feel like they are crawling my recent agent logs to explain and expand the exact thing I’m working on at the time. 😂
So rather than just using Claude, I can now use Claude and Qwen and distill Claude knowledge into a LoRA for Qwen, and then I basically get near the same intelligence as Claude, provided that I keep interacting with the model about the same thing, and I only get this experience after some usage and training. Or ... I could have just used Claude. Sure, Claude is money, and Qwen is free, but even with these training loops, I still have to pay for Claude. If I have a Toyota car, and I want it to run better, it's like I'm being told to buy a Ferrari, take pieces out of the engine of the Ferrari, and put them into the Toyota, and then the Toyota will nearly run as good as the Ferrari. But if I bought a Ferrari, why wouldn't I just drive the Ferrari and ditch the Toyota all together? All this is is perpetual distillation, in which I pay for Claude to distill to Qwen to almost function at a Claude level. Interesting that it works, but I have better thing to waste my time with that provide a higher value of return. Not to mention, what about all the time I use Qwen pre-training, the wait time during traing, the post-trained usage where I'm confident in Qwen's results, yet might be engaging it on a topic it hadn't yet been trained and get a bad answer. None of this is useful.
What are the best Hermes agent memory addons? Is obsidian enough?
Can you check out the StateLM and CaveAgent papers?
It's LLMs all the way down.
More good stuff here. 👏
One of your future videos will be about abstracting and three piece or more structures broken into math.... 🙂
Can't wait to see when 1B parameter model with well trained memory skill will outperform all of the beefy Terra-scale models, that require ton of memory and compute to run. This is really way forward - teach a model how to condense and distill the knowledge from the noise, instead of tuning billions of parameters by treating noise as 'useful data to learn' during training.
More epistemic foraging!