Navigate Select ESC Close

I need to rant about local models

2026-07-07 Science & Technology
57.2k
3.1k
821
Theo - t3․gg
Theo - t3․gg
552.0k subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

GLM 5.2 might be the best open-weight model ever released, but "running it yourself" means 400GB of VRAM, a $75,000 GPU box, and about $2,000 a year in electricity that you didn't account for when trying to build a "free" model for your personal work Thank you General Translation for sponsoring! Check them out at: https://soydev.link/gt Want to sponsor a video? Learn more here: https://soydev.link/sponsor-me Check out my Twitch, Twitter, Discord more at https://t3.gg S/O @Ph4seon3 for the awesome edit 🙏 #ai #coding #programming #localmodels

Top Comments (10)

@skuusonius2660 2026-07-07

Anthropic liked this video so much they extended Fable availability until Sunday. Keep em going Theo

573 2 replies
@w.o.jackson8432 2026-07-07

Didn't realize sending death threats over criticizing local AI was on the table, I gotta step my game up.

549 10 replies
@Shapessoftware 2026-07-07

My favorite AI addicted web dev

256 4 replies
@MarSprite 2026-07-07

I use my local spark model for handling data that cannot be permitted to be run through a cloud provider - because once you put data into a prompt for a cloud model, you have no control over what happens to it. None of these labs can be trusted with private data, they already proved that when they made the models in the first place. So, mostly I use it to read my email, and organize my private data, setting up the process using my frontier AI, but executing it with local inference. Slow ingestion of Air Xiv papers into a database meant for my frontier AI to query. Automations that don't need a strong intelligence.

89 8 replies
@nimya962 2026-07-07

To correct an incomplete point: with MoE models, you can offload experts to RAM, though prefill performance can take a hit—but that's acceptable. I run Qwen3.6-35B-MTP-IMAT-IQ4-XS (which is Q8nextn, sitting at ~19GB) on an RTX 3080 with 8GB VRAM + 32GB system RAM, on a 5-year-old Asus laptop, alongside a small model plus TTS and STT models. Basically, I built my own voice assistant because I like ChatGPT's voice mode—it's just a PWA running over a WireGuard virtual network. I get the audio response in less than 2 seconds most of the time, depending on the request. And yes, it's not flawless, but I use it as a glorified Siri. Besides that, there are breakthroughs in clustering, and I have hope that specialized local AIs will do a better job than SOTA models in their respective fields. Finally, about Zhipu: $40 for 3 months with only a 5-hour limit is insane value without any equivalent. My max usage was 3 billion tokens in a month. So, I'm not a Cassandra. The thing is, PCs weren't built for AI—PCIe speed, RAM, VRAM... There is room for improvement in hardware, architecture, and the software stack. There is room for improvement in the fundamentals of quantization and LLM architecture. There's also the fine-tuning process, which is becoming accessible to everyone to fit their specific usage. BUT local ai is needed in many sectors who for different reason and this could help us making local LLM feasible for coding.

74 10 replies
@hjewkes 2026-07-08

Theo out there burning 25k a month in tokens he's getting comped by cloud providers, and then complains a local LLM would cost $2k a year in electricity

66 10 replies
@AustinGlamourPhoto 2026-07-07

you are thinking of model with a specific use for coding. If you just want a vision caption model, a chatbot, an image generation model, etc. Local models are just fine. You want agentic coding, you need a larger model.

66
@ivanfenenko 2026-07-07

I tried to comment first but my local model just managed to post this reply now

61 1 replies
@ericlippe 2026-07-08

"I am not using one agent, I'm going between zero and forty" You have to realize how insane that cost is to an average person or company. Someone who is tinkering or a company that is not yet profitable can’t afford that long term. Especially if you are running frontier models at their current cost. Are the other options slower? Yeah. Is the hardware insanely priced for open source frontier models? Yes. But a few months of running Opus or GPT5.6 (nevermind Fable or Sol, pending availability) is going to eclipse the hardware cost of a medium-tier GPU setup in a few months of use.

22
@abrahamj15 2026-07-07

COO of an AI company here. Not everyone needs opus 4.6 Small models have reached beyond GPT 4.1 and for customer service, Q&A and context translations have proven a safe option for companies that want to control their data, $ 10k is a small price for a pc that's can run Qwen 3.6 27b at 15 answers per minute at 90% precision.

10

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot