Navigate Select ESC Close

This 744GB Model Shouldn't Fit on Your Laptop. It Does

2026-07-20 Science & Technology
734
61
6
Prompt Engineering
Prompt Engineering
245.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Colibri: Run GLM 5.2 on 25GB of RAM on consumer hardware! A 744B Mixture-of-Experts model activates only ~40B parameters per token — and only ~11 GB of those change from token to token (the routed experts). LINKS: https://github.com/JustVugg/colibri https://z.ai/blog/glm-5.2 DwarfStar-4 Video: https://youtu.be/9gHcmhUDJfw DSpark video: https://youtu.be/eFgknPFK-g0 My voice to text App: whryte.com Website: https://engineerprompt.ai/ RAG Beyond Basics Course: https://prompt-s-site.thinkific.com/courses/rag Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0 Let's Connect: 🦾 Discord: https://discord.com/invite/t4eYQRUcXB ☕ Buy me a Coffee: https://ko-fi.com/promptengineering |🔴 Patreon: https://www.patreon.com/PromptEngineering 💼Consulting: https://calendly.com/engineerprompt/consulting-call 📧 Business Contact: [email protected] Become Member: http://tinyurl.com/y5h28s6h 💻 Pre-configured localGPT VM: https://bit.ly/localGPT (use Code: PromptEngineering for 50% off). Signup for Newsletter, localgpt: https://tally.so/r/3y9bb0

Top Comments (10)

@christopherbrand5360 2026-07-20

Project contributor here. I’m on a Mac M5 Max 128gb and the current main branch defaults get you just over 2tok/s. There are PRs in the queue that get well past 3tok/s, which still feels slow but verging on useable. I am loving the project so far and have some batch (overnight, not interactive) work that I plan to have this beast doing in the near future.

12 1 replies
@12kenbutsuri 2026-07-20

I just tried this on gdx spark. I waited for a couple hours for it to finish one prompt, and after 2 or 3 hours it was pronting "..........." in thinking and hasn't even finished lol

4
@deniskfender 2026-07-20

You can also raid two pcie 5.0 SSDs to have high reading speed. Looks super promising since you can have for example two 16-24gb GPU and keep power consumption under 1kwt.

4
@josephroman2690 2026-07-20

I'm curious how much it would help the speed when using vram, in all the videos available on this platform people just use ram and get around 0.1t/s plus adding mtp how much speed you can get. I'm curious about it.

0
@W1Io_oI1W 2026-07-21

Thank you

0
@alexisdamnit9012 2026-07-21

On the daily I need a language model that contains all of human knowledge within my laptop 😂

0
@stephenbaldwin7165 2026-07-20

Does it scale to 256GB? So today it's too slow. But just a year away and it's totally useable locally! (For £15k machine)

0
@SimplestUsername 2026-07-21

GLM 5.2 is available on DS4 now. I haven't tried it yet so I can't say how it stacks up to Deepseek V4 Flash on a 128gb M5.

0 1 replies
@14supersonic 2026-07-21

The foucs of this shouldn't necessarily be about just running it on storage+RAM, but all 3. Imagine using this on a smaller model like deepseek v4 flash or minimax m3 with vram acceleration and storage+RAM cache. All of a sudden you go from only being able to run the smallest models to having access to practically any model at useful to very usable speeds.

0
@metrodyne 2026-07-21

I really feel MoE sometimes makes the model dumber. If i want to talk about complex matters, i want full parameters, it is not like everyone else are interested in code only, but that's what the market are going for. Language, phylosophy, health, psychology, etc, MoE can be very limited for this.

1

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot