AI Agents Code Their Own Harness Optimization: NO, WAIT!
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Trump's Pentagon PANICS as Their COVER UP is SHUT DOWN!!!
Legal AF
97.2k views
r/AITA I Called the Cops on My Own Parents
rSlash
78.7k views
Michael Jordan Was NEVER A 3rd Option!!
The Arena
78.5k views
IHIP News: 🚨 Dems Vote to HELP Trump's WAR and BETRAY Their OWN Constituents!
I've Had It
170.2k views
Trump Admin FREAKS OUT over their OWN DISASTER
Adam Mockler
403.2k views
🚨Kristi Noem Thrown UNDER THE BUS by…her OWN ICE AGENTS !!!
MeidasTouch
940.9k views
Dirty war continues. NATO wants escalation
The Duran
78.7k views
IHIP News: Trump Gets DIRE WARNING Over the Future of AI! He's OWNED By Tech BILLIONAIRES!
I've Had It
18.3k views
How I code with AI right now
Theo - t3․gg
192.1k views
OpenAI is in "CODE RED" (Did Gemini win that hard??)
Theo - t3․gg
72.0k views
Top Comments (6)
Hoy por hoy este es el mejor canal para quiénes planeen proyectos de software con IA. Inspiración total.
En el caso de cogniteam, no se utiliza un solo Llm para los tres agentes (developer, debbuguer y meta-agente), se utiliza un Llm diferente en cada agente utilizando y optimizando su utilización en capas free, para no sobre pasarlas, reajustando el Llm utilizado en función de este requerimiento.
Harness evolution has to be paired with other elements. For one to beat overfitting ,you need an adversarial red team Judge that can also scale. You have to go from from stochastic self-optimization to adversarial robustness in agent design. This aligns with recent research on autonomous research systems (such as ARIS), which argue that single-model execution is inherently prone to "laziness" or "plausible unsupported success."
I have been doing this for a little over a month. Where I use auto research to optimize my harness configuration around my internal benchmarks. Whenever I have leftover usage on my plans on sunday morning, I burn the rest of it doing this
Yo hice esto alguna vez 😅 no tan así, solo en principio.
Single point of unfillable knowledge. without it all falls apart or to the local minim. Then with a system that breaks the models with single task trained to task. The point is not for a better solution it is for a perfect solution or as close as we can get to it . Yes a fully trained system that has had thousands of hours of training in on the direction of open training does not do well trained to a specific task, go figure, and base ball bats do not work well as knives go figure, and yes asking the system a thousand times rather than once will give better results. That is not what we need. For gods sake am I the only one budling self training systems into the harness. The parts have to all be reflexive if not its like a hermit crab picking shells and let me guess.. none of you changed how it saw information at all. I am working with a self nesting tag system that forms its own tag webs to act as single tokens to train on. with the entire knowledge graph being a dehydrated tag representation. No shit they do bad, and how much did they train obviously not to a point of competence. Which can be tested for hell testing for that would have told them the results from the git. I sware to good I sometimes feel like a person trying to describe a rotating detonation jet engine to people that just understood the piston. Sometimes just sometimes it takes a load of systems to make a dame near impossible thing happen.
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (6)
Hoy por hoy este es el mejor canal para quiénes planeen proyectos de software con IA. Inspiración total.
En el caso de cogniteam, no se utiliza un solo Llm para los tres agentes (developer, debbuguer y meta-agente), se utiliza un Llm diferente en cada agente utilizando y optimizando su utilización en capas free, para no sobre pasarlas, reajustando el Llm utilizado en función de este requerimiento.
Harness evolution has to be paired with other elements. For one to beat overfitting ,you need an adversarial red team Judge that can also scale. You have to go from from stochastic self-optimization to adversarial robustness in agent design. This aligns with recent research on autonomous research systems (such as ARIS), which argue that single-model execution is inherently prone to "laziness" or "plausible unsupported success."
I have been doing this for a little over a month. Where I use auto research to optimize my harness configuration around my internal benchmarks. Whenever I have leftover usage on my plans on sunday morning, I burn the rest of it doing this
Yo hice esto alguna vez 😅 no tan así, solo en principio.
Single point of unfillable knowledge. without it all falls apart or to the local minim. Then with a system that breaks the models with single task trained to task. The point is not for a better solution it is for a perfect solution or as close as we can get to it . Yes a fully trained system that has had thousands of hours of training in on the direction of open training does not do well trained to a specific task, go figure, and base ball bats do not work well as knives go figure, and yes asking the system a thousand times rather than once will give better results. That is not what we need. For gods sake am I the only one budling self training systems into the harness. The parts have to all be reflexive if not its like a hermit crab picking shells and let me guess.. none of you changed how it saw information at all. I am working with a self nesting tag system that forms its own tag webs to act as single tokens to train on. with the entire knowledge graph being a dehydrated tag representation. No shit they do bad, and how much did they train obviously not to a point of competence. Which can be tested for hell testing for that would have told them the results from the git. I sware to good I sometimes feel like a person trying to describe a rotating detonation jet engine to people that just understood the piston. Sometimes just sometimes it takes a load of systems to make a dame near impossible thing happen.