Navigate Select ESC Close

Ornith 1.0: This is new class of self-improving model

2026-06-27 Science & Technology
3.2k
83
21
Prompt Engineering
Prompt Engineering
245.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Ornith 1: Open-Weight Agentic Coding Models That Write Their Own Harnesses (and Beat Bigger Models) In this video: I break down Ornith 1, a new family of open-weight models built for agentic coding that can outperform much larger models on benchmarks like Terminal Bench, with the 397B model nearing closed-source performance (Opus 4.8). The key idea isn’t just scores—it’s how Ornith is trained to generate both solution rollouts and a task-specific harness (memory, retries, error handling) in a single loop, using reinforcement learning (GRPO) so rewards update both the solution and the scaffold. I cover reward-hacking risks and the three-layer defenses: locked boundaries, deterministic monitoring, and a frozen judge model. I also share my own Ollama tests on an M2 Max comparing Qwen 3.5 9B base vs Ornith 1 9B (8-bit): similar accuracy, but Ornith is ~3× cheaper (up to 20× on some tasks), while long-horizon “honesty under pressure” seems to require larger scale (35B+). LINKS: https://deep-reinforce.com/ornith_1_0.html https://huggingface.co/collections/deepreinforce-ai/ornith-10 My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: [email protected] Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 00:00 Ornith Models Overview 02:46 Self Written Harnesses 04:31 Reward Hacking Risks 05:30 Three Layer Defenses 06:17 My Ollama Test Setup 07:40 Private Bench Results 08:33 Long Horizon Honesty Test 09:53 Key Takeaways on 9B 10:42 Caveats

Top Comments (10)

@EarthAaron 2026-06-27

I'd say it's more like self-optimization -- self-improvement would require it improve the actual model, not just the harness -- still pretty cool tho

5 6 replies
@Roberto-x4s1f 2026-06-27

how can a model be self improving? Its read only after training. The agent yes, but i am curious how the model itself can be self improving?

3 5 replies
@rotimiawaye 2026-06-27

This sounds NOT like you

2 3 replies
@MeinDeutschkurs 2026-06-27

BTW: The current excuse for realtime training is that a model degrades constantly through stupid inputs. Is there any real time modifying LLM out there?

1
@akierum 2026-06-27

The Jinja template is wrong, the 31b dense model not public, wtf is the joy about?

1
@weise-25 2026-06-27

9B one is trash unfortunately

0 3 replies
@phoebusapollo82 2026-06-28

Talk louder

0 2 replies
@aryadas1095 2026-06-27

This model is pure trash I used it’s 35b

0
@HassanAllaham 2026-06-28

🌹

0
@arghyashrivastav8559 2026-07-01

Can you test their 35B MoE against 3.6 MoE please Also While they have announced it but they still haven't released the 31B dense version when it does please also compare it against qwen 3.6 27B dense. I am running a 24GB vram setup want to figure out which would be the best choice within constraint. Currently running Qwen 3.6 27B MTP at Q4 quant and Q8KV at a context window of 92k works on agentic tasks better, but trying to figure out how to take context window above 100k or 112k, will try to use MUX switch to off load system to run on IGPU and change LM studio offload completely to GPU then will get access to full 24GB.

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot