This 27B Model Shouldn't Run On Your Phone. It Does.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
This 100% uncensored AI model is insane… let’s run it
David Ondrej
49.6k views
This AI Model Runs On Your Phone (With No Internet)!
Matt Wolfe
24.4k views
Sonnet 4.5 Is Here—And It’s a Beast at Coding
Prompt Engineering
52.0k views
GPT-OSS Jailbreak with this Simple Trick
Prompt Engineering
54.4k views
Context Engineering is All You NEED!
Prompt Engineering
38.7k views
The Only Embedding Model You Need for RAG
Prompt Engineering
35.2k views
AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Lenny's Podcast
68.3k views
THIS RUINED OUR FRIENDSHIP! -You Should Know Podcast- Episode 167
You Should Know Podcast
363.0k views
The Secret to Perfect Prompts (Without Prompt Engineering)
Futurepedia
55.4k views
Do Anything with Local Agents with AnythingLLM
Prompt Engineering
60.4k views
Top Comments (10)
I have noticed the same looping problem with "obliterated" (aka jail broken aka uncensored) models. The trick for me was to just change the inference parameters (like temperature for example) to get it to not do that. This is incredible work!
It was looping on me too But love the concept… hope the field progresses and smaller and smaller models with intelligence density Become common
Microsoft has something like a $13B position in OpenAI and sells inference by the token through Azure. A self-hostable, CPU-runnable, Claude-tier open model is the single most destructive thing to that business
Cool tutorials both on paper explanations and on new model architectures One question might feel unrelated: may I ask what you use for animations within your videos? Feels like manim, but it's not
I find the ternary version quite usable. I haven't had the looping issues you had at this time. However, I also wasn't using the 1-bit version. I think ternary is the way to go if you want to use it for anything useful. One agentic task that seems to work great is a chatbot based on traversing an LLM-Wiki (vs RAG). That seems to be working nearly flawlessly for my LLM-Wiki KBs.
I ran it today in a LXC on my EPYC 7713 and got ~1 token/s per 16 cores, I guess it is not that shiny yet
So how about making smaller - models that can work on yout mac machine? Will this shrink larger models without quality loss?
It's like taking a Ferrari, removing the engine, it's still a Ferrari that may fit in your budget now but you won't go anywhere.
This can be really disruptive if it holds for bigger models.
is there a way to expand the attention so it doesn’t loop?
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
I have noticed the same looping problem with "obliterated" (aka jail broken aka uncensored) models. The trick for me was to just change the inference parameters (like temperature for example) to get it to not do that. This is incredible work!
It was looping on me too But love the concept… hope the field progresses and smaller and smaller models with intelligence density Become common
Microsoft has something like a $13B position in OpenAI and sells inference by the token through Azure. A self-hostable, CPU-runnable, Claude-tier open model is the single most destructive thing to that business
Cool tutorials both on paper explanations and on new model architectures One question might feel unrelated: may I ask what you use for animations within your videos? Feels like manim, but it's not
I find the ternary version quite usable. I haven't had the looping issues you had at this time. However, I also wasn't using the 1-bit version. I think ternary is the way to go if you want to use it for anything useful. One agentic task that seems to work great is a chatbot based on traversing an LLM-Wiki (vs RAG). That seems to be working nearly flawlessly for my LLM-Wiki KBs.
I ran it today in a LXC on my EPYC 7713 and got ~1 token/s per 16 cores, I guess it is not that shiny yet
So how about making smaller - models that can work on yout mac machine? Will this shrink larger models without quality loss?
It's like taking a Ferrari, removing the engine, it's still a Ferrari that may fit in your budget now but you won't go anywhere.
This can be really disruptive if it holds for bigger models.
is there a way to expand the attention so it doesn’t loop?