Navigate Select ESC Close

Local AI models destroyed by further Distillation

2026-07-06 Science & Technology
2.5k
99
14
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

What happens when the perfect AI training method accidentally teaches a model how to cheat? We rely on self-distillation to make small models reason like giant ones: so why do these "student" AIs sometimes lose their minds, pathologically repeating words like 'Wait...' instead of actually thinking? Are our privileged AI teachers silently poisoning their students' tensor weights with hidden shortcuts? Dive into the mathematical paradox that is breaking long-chain-of-thought models, and discover the brilliant algorithmic "filter" researchers just invented to save the AI's ability to reason. All rights w/ authors: Purified OPSD: On-Policy Self-Distillation Without Losing How to Think Zhanming Shen1,2 Jintao Tong2,3 Shaotian Yan2 Chen Shen2†‡ Hao Chen1,2 Wentao Ye1 Xiaomeng Hu1,2 Rui Miao2,4 Haobo Wang1† Junbo Zhao1 Gang Chen1 Jieping Ye2 from 1 Zhejiang University 2 Tongyi Lab, Alibaba Group 3 Huazhong University of Science and Technology 4 Jilin University #airesearch #aitechnology #scienceexplained #aiexplained

Top Comments (10)

@forseasmr 2026-07-06

About the epistemic token count (ETC), the situation is the same as before, so the distillation worked well, reproducing an inefficiency of the model (DeepSeek). Could it be an explanation?

5
@FirstLast-rb5zj 2026-07-06

I am not certain this is such a great discovery. The way these are trained is I believe to keep changing so it improves. It is then increasingly the case that a change makes no difference or makes it worse. They're normally shortened already as much as can be. You could perhaps create one with faster reaction time that is specialised by pruning certain paths if they are overly general. The reasoning though is not real. That's a mistake. It is reproducing the illusion of reasoning but it's not actually thinking or understanding what it is saying like a person would. If you want to improve a model you often need to bloat it which is fiddly, then compact again. If you over train on something specific then a model can like a student just memorise the answer.

4 1 replies
@TheAdeybob 2026-07-06

so...basically, they sussed that the model was shortcutting, and they tried to overcome it. They did so but with limited success - but even limited success might point at complete solutions. As stated, the limited success suggests the solution is incomplete. Without question, the team found a rather good way to preserve the signal...but did they really have enough data from which to draw a strong enough signal in the first place? As a non academic, I'll have to scurry off to an AI for better understanding of this paper and the points made in this vid, but these are my immediate impressions.

3 1 replies
@timmygilbert4102 2026-07-06

Always count on chinese to provide the rigorous mathematical explanation of intuition.🎉

3
@rmt3589 2026-07-06

Simple fix, hook up real intelligence into the loop. Ethical options include human neurons, mycelium, slime mold, ant hive, virtual c eligans(misspelled), etc. Probably unethical solution is a human sacrifice with bci. Only ethical if the owner/CEO of the AI is the willing sacrifice. (Note, sacrifice doesn't mean death here, and they can stop being a sacrifice later, just gotta live with the irreversible changes to their body)

2 2 replies
@tjhazmat2760 2026-07-06

Simple fix. 3:44 You have your valid solution. Teacher generates the synthetic CoT or reasoninv traces to simulate the prompt to valid solution trace. Chunk the CoT or reasoning into components steps. Mask or remove one step of reasoning, have the students learn by recreating the masked out steps.

2 2 replies
@wlatol6512 2026-07-06

So self improving harnesses are obsolete?

1
@konzy2 2026-07-07

I'll take an attempt at explaining the epistemic count and marker distribution graphs. First I'd like to point out, that these are the reasoning tokens they are looking at. When it says, "Wait, maybe I'm overthinking this" or, "Perhaps, we need to think differently." a) Shows the total token count of ["wait", "maybe" "check" and Others] over the training steps. With Qwen, the token count of "wait", "maybe" and "check" drop dramatically. Conversely with Deepseek-R1, token count goes up, nearly doubling "wait" more than doubles. I think what the paper is trying to indicate that this is not an indication of the model learning the result, but instead memorizing through overfitting, because if it was learning, it would make smaller movements from the base/frozen model. This is what we can see in the OPSD-PMI (Ours) bars, smaller movements from the base.

0
@countofserenno7605 2026-07-06

Do one on concolic testing

0
@MirolimMirzakhmedov 2026-07-06

Good paper , thanks for sharing!

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot