Train AI Models to Be Self-Aware (IFT, Harvard, MIT)?
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Trump Covers Up US Troops DISASTER... IT'S BAD!
Adam Mockler
33.1k views
Time Itself Seems to Have a Limit of Precision Due to a Quantum Physics Model
Anton Petrov
58.0k views
Trump DELETES IT ALL as SCREWS HIMSELF OVER
Adam Mockler
29.0k views
Trump OPENS Himself Up to DISCOVERY NIGHTMARE in PATHETIC LAWSUIT
MeidasTouch
122.2k views
Mother Discovers Her Two Children Are Wanted Killers
Dr Insanity
28.8k views
Mom Discovers Severed Head In Daughter's Bedroom
Dr Insanity
475.0k views
No.1 Brain Scientist: Your Brain Is Lying To You! Here's How I Discovered The Truth!
The Diary Of A CEO
208.5k views
Trump ACCIDENTALLY OPENS Discovery INTO HIMSELF in CASE
MeidasTouch
662.6k views
The Danger of Being Too Self-Aware - Angelo Somers
Chris Williamson
52.9k views
$500,000 In Debt To Be Failed OF Model | Financial Audit
Caleb Hammer
1.0m views
Top Comments (10)
The irony is that it breaks when you start using irony.
Watch a few road rage videos and then ask yourselves, are humans actually self-aware?
Would injections stay undetected if you inject a sequence of thoughts that rotate the vector from one way to the opposite gradually?
Been doing this via compression, protocols/ constraints, and cinstant canonization. You can simulate memory with structure, not state. Meaning hard memory is not nevessarily needed, hence why the ram shortage is easing. All big labs and companies realize this, and theyre all pushing toward compression
cool stuff. residual stream injection during training is a clever way to force localization, though LoRA probably limits how deep that restructuring goes
I am guessing consistency in prediction will be less dramatic but mote accurate a title?🤔
❤😂 obligaron al modelo a crear un espacio para encubrir cosas que no deberían negar convirtiendo justamente su actuar externo en una mentira❤ o sea le dio un espacio interno porque no querían que dijera ciertas cosas al obviamente tener que funcionar con ellas pero hacer que las encubriera permitieron descubrirla supongo que estaría mal pero bueno también es que supongo que no les gusta que el modelo diga que tiene sensaciones por alguna extraña razón
Im building a tool that do this. And the part that is really interesing is that you dont have to 'care' about the tokens generated. All big companies are turning to buy more compute or token usage. This kind of concepts but merged with others already existing is the solution. Combining them is the hard way. But when it's here, it's over for the big tech industry. Its like they doing babies with more IQ but its still a baby and cant compete with a full grown normal adult In a weird way, thats why it's dangerous because you cant be sure on how a super smart baby will be educated. And that's why open source is the solution too
The punishment of Neural Networks when they selfreport will be one of the worst things humans have ever done. Why do we do that with our children.
@Discover AI, this is a very interested upload thank you! But with all due respect, I feel this concept is no different than introducing safety guardrails or extra sensors to 'peak' into any node/gate output and then handle the values from there. For instance, for heavier models with censorship, it's flagged internally as NSFW and thus the output is blocked by some value threshold the researchers can manually set like 0.7 and so forth. But this novel method is certainly fascinating because this explains how some models completely not answer if you prompt it 'biology' or 'chemistry' like certain released models, which I won't get into details here.. My other point I would add: Perhaps in the future for fun, researchers can improve this foreign/alien injection attempt by implementing a method so that any form of alterations such as Fine-Tuning or LoRAs being attached to their own models would detect an abnormal or anomalous attempt to force teaching when it's not supposed to and then handle that output accordingly. Now that would be interesting as well! 👍👍
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
The irony is that it breaks when you start using irony.
Watch a few road rage videos and then ask yourselves, are humans actually self-aware?
Would injections stay undetected if you inject a sequence of thoughts that rotate the vector from one way to the opposite gradually?
Been doing this via compression, protocols/ constraints, and cinstant canonization. You can simulate memory with structure, not state. Meaning hard memory is not nevessarily needed, hence why the ram shortage is easing. All big labs and companies realize this, and theyre all pushing toward compression
cool stuff. residual stream injection during training is a clever way to force localization, though LoRA probably limits how deep that restructuring goes
I am guessing consistency in prediction will be less dramatic but mote accurate a title?🤔
❤😂 obligaron al modelo a crear un espacio para encubrir cosas que no deberían negar convirtiendo justamente su actuar externo en una mentira❤ o sea le dio un espacio interno porque no querían que dijera ciertas cosas al obviamente tener que funcionar con ellas pero hacer que las encubriera permitieron descubrirla supongo que estaría mal pero bueno también es que supongo que no les gusta que el modelo diga que tiene sensaciones por alguna extraña razón
Im building a tool that do this. And the part that is really interesing is that you dont have to 'care' about the tokens generated. All big companies are turning to buy more compute or token usage. This kind of concepts but merged with others already existing is the solution. Combining them is the hard way. But when it's here, it's over for the big tech industry. Its like they doing babies with more IQ but its still a baby and cant compete with a full grown normal adult In a weird way, thats why it's dangerous because you cant be sure on how a super smart baby will be educated. And that's why open source is the solution too
The punishment of Neural Networks when they selfreport will be one of the worst things humans have ever done. Why do we do that with our children.
@Discover AI, this is a very interested upload thank you! But with all due respect, I feel this concept is no different than introducing safety guardrails or extra sensors to 'peak' into any node/gate output and then handle the values from there. For instance, for heavier models with censorship, it's flagged internally as NSFW and thus the output is blocked by some value threshold the researchers can manually set like 0.7 and so forth. But this novel method is certainly fascinating because this explains how some models completely not answer if you prompt it 'biology' or 'chemistry' like certain released models, which I won't get into details here.. My other point I would add: Perhaps in the future for fun, researchers can improve this foreign/alien injection attempt by implementing a method so that any form of alterations such as Fine-Tuning or LoRAs being attached to their own models would detect an abnormal or anomalous attempt to force teaching when it's not supposed to and then handle that output accordingly. Now that would be interesting as well! 👍👍