Navigate Select ESC Close

OpenThoughts: The Blueprint for Agentic SFT

2026-06-26 Science & Technology
1.2k
63
6
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Everyone talks about building better AI agents by scaling models or adding smarter tools or harnesses. But what if the real bottleneck is the training data itself? In this video, I break down OpenThoughts-Agent, a new research project that doesn’t just train an agent: it systematically engineers the best possible supervised fine-tuning dataset for long-horizon coding and terminal agents. We’ll go through how the authors ablate task sources, teacher models, rollout quality, and dataset filtering to uncover what actually makes an agent learn to act. If you want to understand where the next generation of open AI agents will come from, this pre-print is far more important than it first looks. All rights w/ authors: OPEN THOUGHTS AGENTS Data Recipes for Agentic Models Negin Raoof∗1, Richard Zhuang∗2, Marianna Nezhurina∗3,4,5, Etash Guha∗2, Atula Tejaswi6, Ryan Marten7, Charlie F. Ruan1, Tyler Griggs1, Alexander Glenn Shaw8, Hritik Bansal9, E. Kelly Buchanan2, Artem Gazizov10, Reinhard Heckel11, Chinmay Hegde12, Sankalp Jajee13, Daanish Khazi14, Emmanouil Koukoumidis15, Xiangyi Li16, Hange Liu17, Shlok Natarajan2, Harsh Raj18, Nicholas Roberts19, Ethan Shen20, Nishad Singhi21, Michael Siu22, Ashima Suvarna9, Hanwen Xing22, Patrick Yubeaton12, Robert Zhang6, Leon Liangyu Chen2, Xiaokun Chen2, Steven Dillmann2, Saadia Gabriel9, Xunyi Jiang23, Anurag Kashyap24, Boxuan Li25, Yein Park26, Minh Pham12, Sujay Sanghavi6, Lin Shi27, Ke Sun17, Yixin Wang28, Zhiwei Xu28, Erica Zhang2, Siyan Zhao9, Wanjia Zhao2, Jenia Jitsev3,4,5, Alex Dimakis1,7, Benjamin Feuer†2,15, Ludwig Schmidt†2 from 1UC Berkeley, 2Stanford University, 3JSC, 4LAION, 5Open-Ψ (Open-Sci) Collective, 6University of Texas at Austin, 7Bespoke Labs, 8Laude Institute, 9UCLA, 10Harvard University & Harvard Medical School, 11TU Munich & Munich Center for Machine Learning, 12New York University, 13Medical University of South Carolina, 14The LLM Data Company, 15Oumi.AI, 16BenchFlow, 17Independent Researcher, 18Northeastern University, 19University of Wisconsin–Madison, 20University of Washington, 21TU Darmstadt, 22University of Southern California, 23UC San Diego, 24Amazon, 25Microsoft, 26Korea University, 27Cornell Tech, 28University of Michigan #airesearch #openai #aitechnology #aiagents #aiexplained

Top Comments (9)

@dailydj5555 2026-06-26

I took a glance at their work and saw a few glaring issues that have always made improvements in my training sets. There needs to be noisy examples, mistakes, spelling issues, etc., in the prompts. I would add a 7th stage that only modified prompt examples (not responses) and added noise to them. It has always increased problem solving and creativity in the model trained on that data. I see some other issues past that, too.

2 3 replies
@terryhernandez2768 2026-06-29

HARVARD WAS AMONG THE FIRST TO LOSE THEIR THEOLOGICAL CITING CAPABILITIES, WHY WOULD ANYONE TRUST THEM?

1 1 replies
@AbraOttoMev 2026-06-26

Thank You! 🙏🏾

1
@AlbinoMarine 2026-06-26

RL adding marginal returns on 8B models matches what I see in robotic manipulation. Sample efficiency remains a fundamental bottleneck. Well-curated SFT data often beats adding training complexity.

1
@ChronosProgrammatica 2026-06-27

@Discover AI, With all due respect, although this paper was recent, to use Qwen3.0 at 32B is already outdated! As now we have Qwen 3.6 32B models and 35B with A3B and so forth. I suppose the general idea is the same, but to run Supervised Fine Tuning under these models will probably result in even more robust and stronger results. I understand having dozens of Universities involved takes a significant amount of time and coordination yes, but we're entering a special period where these AI models are now being released at breakneck speeds that even the hardest, hardcore of AI youtube content creators have difficulty catching up. 😔😔

0 3 replies
@timmygilbert4102 2026-06-26

Basically: - surface diversity -> better parsing - trajectory richness -> macro state recovery

0 2 replies
@bjmay67 2026-06-26

I wonder in your paper filtering if you can sort method papers by claimed magnitude of improvement. It seems like almost all of these methods (post-training, training data, skill files, memory schemes, harness loops, etc) give marginal Pareto improvements, which is *expected* if done on top of a base models to squeeze a bit more juice out of it. But are there any papers claim >+50% relative or absolute performance across the board (w/o catastrophically degrading other performance)?

0 2 replies
@benjamindemontgomery6317 2026-06-27

so much drama about AI on You Tube now. this is nice and knowledge booster.

0 1 replies
@QuantumPrintum 2026-07-03

filtering for longer trajectories + synthetic augmentation as the big levers... data quality thesis keeps getting validated. 100 ablations to confirm what most people skip

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot