This 744GB Model Shouldn't Fit on Your Laptop. It Does
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Sonnet 4.5 Is Here—And It’s a Beast at Coding
Prompt Engineering
52.0k views
GPT-OSS Jailbreak with this Simple Trick
Prompt Engineering
54.4k views
Context Engineering is All You NEED!
Prompt Engineering
38.7k views
The Only Embedding Model You Need for RAG
Prompt Engineering
35.2k views
Gemini CLI — Google’s Free Open-Source Coding Agent
Prompt Engineering
56.6k views
AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Lenny's Podcast
68.3k views
The Secret to Perfect Prompts (Without Prompt Engineering)
Futurepedia
55.4k views
Do Anything with Local Agents with AnythingLLM
Prompt Engineering
60.4k views
LightRAG: A More Efficient Solution than GraphRAG for RAG Systems?
Prompt Engineering
84.5k views
EASIEST Way to Fine-Tune LLAMA-3.2 and Run it in Ollama
Prompt Engineering
100.8k views
Top Comments (10)
Project contributor here. I’m on a Mac M5 Max 128gb and the current main branch defaults get you just over 2tok/s. There are PRs in the queue that get well past 3tok/s, which still feels slow but verging on useable. I am loving the project so far and have some batch (overnight, not interactive) work that I plan to have this beast doing in the near future.
I just tried this on gdx spark. I waited for a couple hours for it to finish one prompt, and after 2 or 3 hours it was pronting "..........." in thinking and hasn't even finished lol
You can also raid two pcie 5.0 SSDs to have high reading speed. Looks super promising since you can have for example two 16-24gb GPU and keep power consumption under 1kwt.
I'm curious how much it would help the speed when using vram, in all the videos available on this platform people just use ram and get around 0.1t/s plus adding mtp how much speed you can get. I'm curious about it.
Thank you
On the daily I need a language model that contains all of human knowledge within my laptop 😂
Does it scale to 256GB? So today it's too slow. But just a year away and it's totally useable locally! (For £15k machine)
GLM 5.2 is available on DS4 now. I haven't tried it yet so I can't say how it stacks up to Deepseek V4 Flash on a 128gb M5.
The foucs of this shouldn't necessarily be about just running it on storage+RAM, but all 3. Imagine using this on a smaller model like deepseek v4 flash or minimax m3 with vram acceleration and storage+RAM cache. All of a sudden you go from only being able to run the smallest models to having access to practically any model at useful to very usable speeds.
I really feel MoE sometimes makes the model dumber. If i want to talk about complex matters, i want full parameters, it is not like everyone else are interested in code only, but that's what the market are going for. Language, phylosophy, health, psychology, etc, MoE can be very limited for this.
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
Project contributor here. I’m on a Mac M5 Max 128gb and the current main branch defaults get you just over 2tok/s. There are PRs in the queue that get well past 3tok/s, which still feels slow but verging on useable. I am loving the project so far and have some batch (overnight, not interactive) work that I plan to have this beast doing in the near future.
I just tried this on gdx spark. I waited for a couple hours for it to finish one prompt, and after 2 or 3 hours it was pronting "..........." in thinking and hasn't even finished lol
You can also raid two pcie 5.0 SSDs to have high reading speed. Looks super promising since you can have for example two 16-24gb GPU and keep power consumption under 1kwt.
I'm curious how much it would help the speed when using vram, in all the videos available on this platform people just use ram and get around 0.1t/s plus adding mtp how much speed you can get. I'm curious about it.
Thank you
On the daily I need a language model that contains all of human knowledge within my laptop 😂
Does it scale to 256GB? So today it's too slow. But just a year away and it's totally useable locally! (For £15k machine)
GLM 5.2 is available on DS4 now. I haven't tried it yet so I can't say how it stacks up to Deepseek V4 Flash on a 128gb M5.
The foucs of this shouldn't necessarily be about just running it on storage+RAM, but all 3. Imagine using this on a smaller model like deepseek v4 flash or minimax m3 with vram acceleration and storage+RAM cache. All of a sudden you go from only being able to run the smallest models to having access to practically any model at useful to very usable speeds.
I really feel MoE sometimes makes the model dumber. If i want to talk about complex matters, i want full parameters, it is not like everyone else are interested in code only, but that's what the market are going for. Language, phylosophy, health, psychology, etc, MoE can be very limited for this.