OpenThoughts: The Blueprint for Agentic SFT
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
This Is The 2028 Blueprint For Democrats
The Majority Report w/ Sam Seder
12.5k views
2026 Discovery: The Platypus is Even Weirder Than We Thought
Anton Petrov
74.2k views
Open Source might change forever
The PrimeTime
68.2k views
Mark Cuban’s Blueprint for Affordable Healthcare
The Bulwark
71.9k views
OpenAI’s Sam Altman ALREADY Looking For A Government Bailout!
The Jimmy Dore Show
61.8k views
Trump Henchman's Blueprint To Dismantle The Federal Government
The Majority Report w/ Sam Seder
28.2k views
Zohran's Blueprint For Beating The Establishment
The Majority Report w/ Sam Seder
58.6k views
The Map Is the Weapon. Roland and Cliff Albright Expose the Blueprint for Black Voter Erasure.
Roland S. Martin
31.9k views
The Blueprint for Fighting Back. HBCU Prez $1M Lawsuit Puts Racism on Trial.
Roland S. Martin
38.7k views
They FOUGHT Amazon’s $3.6B AI Data Center
Breaking Points
43.5k views
Top Comments (9)
I took a glance at their work and saw a few glaring issues that have always made improvements in my training sets. There needs to be noisy examples, mistakes, spelling issues, etc., in the prompts. I would add a 7th stage that only modified prompt examples (not responses) and added noise to them. It has always increased problem solving and creativity in the model trained on that data. I see some other issues past that, too.
HARVARD WAS AMONG THE FIRST TO LOSE THEIR THEOLOGICAL CITING CAPABILITIES, WHY WOULD ANYONE TRUST THEM?
Thank You! 🙏🏾
RL adding marginal returns on 8B models matches what I see in robotic manipulation. Sample efficiency remains a fundamental bottleneck. Well-curated SFT data often beats adding training complexity.
@Discover AI, With all due respect, although this paper was recent, to use Qwen3.0 at 32B is already outdated! As now we have Qwen 3.6 32B models and 35B with A3B and so forth. I suppose the general idea is the same, but to run Supervised Fine Tuning under these models will probably result in even more robust and stronger results. I understand having dozens of Universities involved takes a significant amount of time and coordination yes, but we're entering a special period where these AI models are now being released at breakneck speeds that even the hardest, hardcore of AI youtube content creators have difficulty catching up. 😔😔
Basically: - surface diversity -> better parsing - trajectory richness -> macro state recovery
I wonder in your paper filtering if you can sort method papers by claimed magnitude of improvement. It seems like almost all of these methods (post-training, training data, skill files, memory schemes, harness loops, etc) give marginal Pareto improvements, which is *expected* if done on top of a base models to squeeze a bit more juice out of it. But are there any papers claim >+50% relative or absolute performance across the board (w/o catastrophically degrading other performance)?
so much drama about AI on You Tube now. this is nice and knowledge booster.
filtering for longer trajectories + synthetic augmentation as the big levers... data quality thesis keeps getting validated. 100 ablations to confirm what most people skip
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (9)
I took a glance at their work and saw a few glaring issues that have always made improvements in my training sets. There needs to be noisy examples, mistakes, spelling issues, etc., in the prompts. I would add a 7th stage that only modified prompt examples (not responses) and added noise to them. It has always increased problem solving and creativity in the model trained on that data. I see some other issues past that, too.
HARVARD WAS AMONG THE FIRST TO LOSE THEIR THEOLOGICAL CITING CAPABILITIES, WHY WOULD ANYONE TRUST THEM?
Thank You! 🙏🏾
RL adding marginal returns on 8B models matches what I see in robotic manipulation. Sample efficiency remains a fundamental bottleneck. Well-curated SFT data often beats adding training complexity.
@Discover AI, With all due respect, although this paper was recent, to use Qwen3.0 at 32B is already outdated! As now we have Qwen 3.6 32B models and 35B with A3B and so forth. I suppose the general idea is the same, but to run Supervised Fine Tuning under these models will probably result in even more robust and stronger results. I understand having dozens of Universities involved takes a significant amount of time and coordination yes, but we're entering a special period where these AI models are now being released at breakneck speeds that even the hardest, hardcore of AI youtube content creators have difficulty catching up. 😔😔
Basically: - surface diversity -> better parsing - trajectory richness -> macro state recovery
I wonder in your paper filtering if you can sort method papers by claimed magnitude of improvement. It seems like almost all of these methods (post-training, training data, skill files, memory schemes, harness loops, etc) give marginal Pareto improvements, which is *expected* if done on top of a base models to squeeze a bit more juice out of it. But are there any papers claim >+50% relative or absolute performance across the board (w/o catastrophically degrading other performance)?
so much drama about AI on You Tube now. this is nice and knowledge booster.
filtering for longer trajectories + synthetic augmentation as the big levers... data quality thesis keeps getting validated. 100 ablations to confirm what most people skip