Microsoft Just Dropped LLM's Frontier Data Engineering Secrets
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Michael Cohen’s SECRET MEETING With Trump JUST LEAKED!
Luke Beasley
32.4k views
Trump’s SECRET DRUG just got LEAKED
Adam Mockler
427.2k views
Matt Pocock’s Agentic Engineering Workflow (just copy him)
David Ondrej
58.6k views
A Year Into Making LLMs, and now Topped Open Source SoTA?!
bycloud
22.1k views
Microsoft JUST BROKE OpenAI...
Wes Roth
36.2k views
Rogan Just OUTED TRUMP’S DAMNING SECRET!
Luke Beasley
24.8k views
JD Vance SCARY Epstein BOMB Just DROPPED...
Adam Mockler
172.6k views
A new way to fine-tune LLMs just dropped
bycloud
16.3k views
Engineering The Perfect First Date
Mark Rober
209.5k views
Why can’t LLMs just LEARN the context window?
bycloud
30.9k views
Top Comments (10)
"Microsoft is a black hole of money and talent" - some guy
10:45 I do believe that Deepseek engineers have an extreme drive for optimization. It is probably some small amount of epople that really like to optimize stuff...
Great move from them ! I was wondering how I would go about making a big LLM and couldn't find good sources about all of this
I mean regarding microsoft locking in, did you hear about the Xbox layoffs, they want to restructure the company because at some teams there were 14 Layers of hierarchy - that is reflective of whole microsoft, so no wonder why cant build good product consistently
A loose thought that came to mind: will we see fabricated research to slow down competitors. did it maybe already happen?
check out my project https://www.intuitiveai.academy/ to learn technical LLMs intuitively and you can use limited time code "LOCKIN" for 35% off yearly plan!
The switching between MoE and Dense might have been effective for training in their data center, but I am doubtful it will be as good for inference. The data handling had some really cool ideas. It avoids risk of locking in bad thinking patterns learned by earlier models, and gives a good base if you want to including synthetic data. We saw how hard it can be to remove bad patterns once they start self reinforcing, with the em dashes and goblins.
Ai2 has been awesome, learned a lot from olmo and their papers. Didn't know about the strategy change, thanks for the info
Garbage in = garbage out, always a good reminder, and this paper is a great showcase how clean curated dataset improves model scaling. The architecture looks beautiful, too.
The point about small-scale experiments not always predicting large-scale behavior was fascinating. It’s a good reminder that in frontier AI, scaling isn't just "more of the same", entirely new behaviors and trade-offs can emerge. Really insightful breakdown.
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
"Microsoft is a black hole of money and talent" - some guy
10:45 I do believe that Deepseek engineers have an extreme drive for optimization. It is probably some small amount of epople that really like to optimize stuff...
Great move from them ! I was wondering how I would go about making a big LLM and couldn't find good sources about all of this
I mean regarding microsoft locking in, did you hear about the Xbox layoffs, they want to restructure the company because at some teams there were 14 Layers of hierarchy - that is reflective of whole microsoft, so no wonder why cant build good product consistently
A loose thought that came to mind: will we see fabricated research to slow down competitors. did it maybe already happen?
check out my project https://www.intuitiveai.academy/ to learn technical LLMs intuitively and you can use limited time code "LOCKIN" for 35% off yearly plan!
The switching between MoE and Dense might have been effective for training in their data center, but I am doubtful it will be as good for inference. The data handling had some really cool ideas. It avoids risk of locking in bad thinking patterns learned by earlier models, and gives a good base if you want to including synthetic data. We saw how hard it can be to remove bad patterns once they start self reinforcing, with the em dashes and goblins.
Ai2 has been awesome, learned a lot from olmo and their papers. Didn't know about the strategy change, thanks for the info
Garbage in = garbage out, always a good reminder, and this paper is a great showcase how clean curated dataset improves model scaling. The architecture looks beautiful, too.
The point about small-scale experiments not always predicting large-scale behavior was fascinating. It’s a good reminder that in frontier AI, scaling isn't just "more of the same", entirely new behaviors and trade-offs can emerge. Really insightful breakdown.