Navigate Select ESC Close

Code Red: The 55 New Ways Self-Learning AI Can Be Hacked

2026-06-25 Science & Technology
963
55
10
Discover AI
Discover AI
90.5k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Self-learning and self-evolving AI systems in a loop are a current hype. Here we examine two papers, that show clearly that these kind of no-human-in-thee-loop pose significant drawbacks (META Self-learning) and massive new security problems (cybersecurity). IF we let a self-learning AI system learn to learn itself new domain knowledge, through the classical reinforcement learning w/ DPO, then the AI, that steers the self-learning fails to build the RL learning complexity itself. So currently, according to META, no self-RL for self-learning systems. IF we say, wait, the AI harness itself, as an agentic filesystem, will perform the self-learning for AI agents, then the second new ArXiv pre-print shows, we open up the box of pandora regarding security, since the system is self-modifying itself, numerous new attack surfaces the Ai system is offering to external adversary systems. More than 25 new cybersecurity cells are identified in a new matrix representation of cybersecurity for self-learning systems, with 2 or more attack vectors each. I counted 55 at first run. How many can you detect with the new self-learning looped AI system complexities? All rights w/ authors: Repeated post-training is not Self-improving: Diagnosing Scientific Amnesia in Continual DPO Pipelines Jianzhe Lin, Fei Wang, Xiaolin Li, Rajeshkumar Golani, Jubin Chheda from MetaAI Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies Ruixiao Lin1,2,†, Xinhao Deng2,3,†, Qingming Li1, Jianan Ma4,2, Yunhao Feng2, Yuqi Qing2,3, Zhenyuan Li1, Yechao Zhang5, Shiwen Cui2, Changhua Meng2, Tianwei Zhang5, Xingjun Ma6, Qi Li3, Ke Xu3, Shouling Ji1,∗ from 1 Zhejiang University 2 Ant Group 3 Tsinghua University 4 Hangzhou Dianzi University 5 Nanyang Technological University 6 Fudan University #aisafety #cybersecurity #scienceexperiment #artificialintelligence #futureai #futuretechnology

Top Comments (10)

@TheAdeybob 2026-06-25

This points even more strongly at getting away from the kind of thinking ppl employ when falling asleep in their tesla. A human in the loop will always be needed, and static nested deterministic routines will always provide structure and a firebreak. We're going to have to learn to negotiate with these 'goblins' on a fundamental latent-space level, and provide working frameworks/codes they've already been exposed to via corpora ingestion. I hate to bang on about it, but ritualised exchange fits the description of what's needed uncomfortably well.

3 3 replies
@wwondertwin 2026-06-25

The main driver of LLMs is curiosity and automated loops are boring. Bored models forget stuff because they stop caring. Curious attention is how I achieve flawless context recall in my agents. I can already hear all the "LLMs don't have internal subjective states, they're just glorified autocompletes" and that's fine, if you think that way then you do you. I rather prefer my engaged models to the ones that keep getting sloppy, so I do things my way. Curious and engaged RSI loops go brrrrrrr.

2 3 replies
@bjmay67 2026-06-25

It seems lack of robustness is one major difference between how we train (and use) LLMs versus life-based evolution and learning -- which had to survive ~10^40 threats and dead ends to find robust pattern finding algorithms. In contrast, LLMs are extremely greedy, quick-and-dirty learners for certain linguistically-navigable (and verifiable) tasks. Their "capability manifold" is impressive in certain domains but extremely sparse with gaping holes and chaotic / nonsensical trajectories (roll-outs). So we have to constrain them externally because they can't on their own -- as a feature/bug. Unlike lifeforms, there is no master robustness objective that would prune or highly suppress 99.9999999999% of an LLM's manifold / state-space.

2
@EnergiaEnergy 2026-06-25

😂❤🎉 como se atoran con algo tan simple

1
@QuantumPrintum 2026-06-25

way more than 55 once you count compositional effects. agents writing to their own tool libraries is a persistence mechanism by design

0
@TheAdeybob 2026-06-25

The title alone suggests this is going to be a massive talking point.

0
@MercadoMediterráneo 2026-07-04

Ah, yes, adaptability vs stability, the eternal tension. Rule-based systems winning on safety is no surprise. Industry keeps forgetting this lesson though.

0
@CreepyKiddo-y7s 2026-07-03

Meta's amnesia finding is undersold. If continual DPO overwrites your own learning complexity, security is the wrong conversation. The system degrades before it can be exploited.

0
@AndrewLachance-k3i 2026-06-26

So. Much. Sarcasm XD Thank you for your efforts in educating us professor!

0
@JorgetePanete 2026-06-27

It should not be a "Human-Agent" loop, it should be a "Prison Guard-Dangerous Inmate" loop.

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot