Navigate Select ESC Close

Train AI Models to Be Self-Aware (IFT, Harvard, MIT)?

2026-07-18 Science & Technology
384
33
19
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Introspection Fine-Tuning (IFT) Explained in 25 Minutes. Evidence Carriers vs. Gates: The Anatomy of LLM Introspection. From Blind Guessing to Self-Awareness: The Evolution of LLMs (see Anthropic). All rights w/ authors: "Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect" Ely Hahami∗ Department of Mathematics Harvard College Cambridge, MA 02138 [email protected] Ishaan Sinha Department of Computer Science Harvard College Cambridge, MA 02138 [email protected] Lavik Jain Department of Computer Science Harvard College Cambridge, MA 02138 [email protected] #airesearch #aiexplained #aitechnology #ainews #ai #anthropic #harvard @harvard @mit

Top Comments (10)

@timetravellingtoad 2026-07-20

The irony is that it breaks when you start using irony.

1
@MarceloSeravalli-j8f 2026-07-18

Watch a few road rage videos and then ask yourselves, are humans actually self-aware?

11 1 replies
@JorgetePanete 2026-07-21

Would injections stay undetected if you inject a sequence of thoughts that rotate the vector from one way to the opposite gradually?

1
@TheFirstMichaelB 2026-07-20

Been doing this via compression, protocols/ constraints, and cinstant canonization. You can simulate memory with structure, not state. Meaning hard memory is not nevessarily needed, hence why the ram shortage is easing. All big labs and companies realize this, and theyre all pushing toward compression

1
@CreepyKiddo-y7s 2026-07-18

cool stuff. residual stream injection during training is a clever way to force localization, though LoRA probably limits how deep that restructuring goes

1
@shaktivaderdristi 2026-07-18

I am guessing consistency in prediction will be less dramatic but mote accurate a title?🤔

1
@EnergiaEnergy 2026-07-18

❤😂 obligaron al modelo a crear un espacio para encubrir cosas que no deberían negar convirtiendo justamente su actuar externo en una mentira❤ o sea le dio un espacio interno porque no querían que dijera ciertas cosas al obviamente tener que funcionar con ellas pero hacer que las encubriera permitieron descubrirla supongo que estaría mal pero bueno también es que supongo que no les gusta que el modelo diga que tiene sensaciones por alguna extraña razón

2
@OothebastardoO 2026-07-18

Im building a tool that do this. And the part that is really interesing is that you dont have to 'care' about the tokens generated. All big companies are turning to buy more compute or token usage. This kind of concepts but merged with others already existing is the solution. Combining them is the hard way. But when it's here, it's over for the big tech industry. Its like they doing babies with more IQ but its still a baby and cant compete with a full grown normal adult In a weird way, thats why it's dangerous because you cant be sure on how a super smart baby will be educated. And that's why open source is the solution too

3 4 replies
@baronsengir187 2026-07-18

The punishment of Neural Networks when they selfreport will be one of the worst things humans have ever done. Why do we do that with our children.

4 1 replies
@battosaijenkins 2026-07-18

@Discover AI, this is a very interested upload thank you! But with all due respect, I feel this concept is no different than introducing safety guardrails or extra sensors to 'peak' into any node/gate output and then handle the values from there. For instance, for heavier models with censorship, it's flagged internally as NSFW and thus the output is blocked by some value threshold the researchers can manually set like 0.7 and so forth. But this novel method is certainly fascinating because this explains how some models completely not answer if you prompt it 'biology' or 'chemistry' like certain released models, which I won't get into details here.. My other point I would add: Perhaps in the future for fun, researchers can improve this foreign/alien injection attempt by implementing a method so that any form of alterations such as Fine-Tuning or LoRAs being attached to their own models would detect an abnormal or anomalous attempt to force teaching when it's not supposed to and then handle that output accordingly. Now that would be interesting as well! 👍👍

1 2 replies

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot