"Just Go Local" Solves Nothing. Here's What Does
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Unlock all features
FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.
Related videos
Software engineering is dead now
Theo - t3․gg
165.4k views
Nothing to see here, just chaos everywhere
David Pakman Show
19.3k views
Sonnet 4.5 Is Here—And It’s a Beast at Coding
Prompt Engineering
52.0k views
GPT-OSS Jailbreak with this Simple Trick
Prompt Engineering
54.4k views
What is AI Engineering
Telusko
67.6k views
Did AI Just Kill Software Devs? ex-Google VP of Engineering Speaks out...
Wes Roth
45.8k views
Context Engineering is All You NEED!
Prompt Engineering
38.7k views
Gemini CLI — Google’s Free Open-Source Coding Agent
Prompt Engineering
56.6k views
AI prompt engineering in 2025: What works and what doesn’t | Sander Schulhoff
Lenny's Podcast
68.3k views
Tony Heller: What Does the Actual Data, Geology and Engineering Say??
Ivor Cummins
17.1k views
Top Comments (10)
Running local models is not really about cost, but privacy. In some areas it is just not worthing taking the risk of uploading sensitive data to inference provider.
Huawei is selling Atlas 300i GPUs with 96GB VRAM on consumer market for $2000. If this continues, consumers will have access to hardware capable of running large models, especially if hardware is sold as a "it just works" distribution model.
The problem I have with the "theory" here is it assumes facts not in evidence. 1.) it assumes Frontier models are now, and will always be superior to Apache 2/MIT licensed models. Nothing could be further from the truth. Most Frontier models have become _worse_ in "real-world" functionality lately. 2.) It assumes on-premise compute hardware cost will always be ridiculously inflated due to the "Circular Funding financial fraud" that is today's only AI business model. 3.) It assumes businesses actually need the full AI suite of capabilities. Newsflash: The vast majority of business in the world are NOT software companies. These businesses do need the enhanced automation capabilities that AI already has...even the Apache 2/MIT licensed models. 4.) It assumes local on-site models must have the same T/s performance as Frontier models, and thus every business must own a massive datacenter filled with billions of GPU's. A Strix-Halo or Mac small multi-node cluster can handle Deepseek V4 Pro FP4/INT4 (Although no entity likely requires that full model) - and certainly ds4 flash/Q4 on a pair of unified memory nodes. I do agree with the escalation scenario presented, but the other _Required_ aspect of that is a local semantic caching library specifically designed to store validated Frontier LLM responses so the local model never needs to escalate for that again. Individuals and 99.9% of all businesses do not need to solve for all the ills of the world, they just need their fairly static small needs met. Thus Hyperscaler Frontier models will become less and less important for smart people that pursue a progressive adoption of local AI.
I think the future is probably local LLMs, but with some connection to superior models online for more complex requests. Maybe even renting computing power to them.
The problem is deciding which sub-task is easy without actually solving it.
I've been benchmarking models against 4,000 very subtle multiple choice science questions. gemma-4-26b-a4b-it is effectively as good as any other model I've tested so far (within a few percentage points), including Opus and Gemini Pro.
a router that can intercept the prompt in chat? as a token saving process
so how about how about xiaomi mimo and it's licensing?
all depends on training.
i keep waiting for the day a full state of the art model can be run in 1gb of ram 😅
Unlock the Data Inside
Turn Videos into Knowledge
- Get FREE 10/day: transcripts, summaries, chats
- Chat with videos, export text & PDF
- $1 free API credit for RAG, chatbots & research
Free forever plan • All features unlocked
Top Comments (10)
Running local models is not really about cost, but privacy. In some areas it is just not worthing taking the risk of uploading sensitive data to inference provider.
Huawei is selling Atlas 300i GPUs with 96GB VRAM on consumer market for $2000. If this continues, consumers will have access to hardware capable of running large models, especially if hardware is sold as a "it just works" distribution model.
The problem I have with the "theory" here is it assumes facts not in evidence. 1.) it assumes Frontier models are now, and will always be superior to Apache 2/MIT licensed models. Nothing could be further from the truth. Most Frontier models have become _worse_ in "real-world" functionality lately. 2.) It assumes on-premise compute hardware cost will always be ridiculously inflated due to the "Circular Funding financial fraud" that is today's only AI business model. 3.) It assumes businesses actually need the full AI suite of capabilities. Newsflash: The vast majority of business in the world are NOT software companies. These businesses do need the enhanced automation capabilities that AI already has...even the Apache 2/MIT licensed models. 4.) It assumes local on-site models must have the same T/s performance as Frontier models, and thus every business must own a massive datacenter filled with billions of GPU's. A Strix-Halo or Mac small multi-node cluster can handle Deepseek V4 Pro FP4/INT4 (Although no entity likely requires that full model) - and certainly ds4 flash/Q4 on a pair of unified memory nodes. I do agree with the escalation scenario presented, but the other _Required_ aspect of that is a local semantic caching library specifically designed to store validated Frontier LLM responses so the local model never needs to escalate for that again. Individuals and 99.9% of all businesses do not need to solve for all the ills of the world, they just need their fairly static small needs met. Thus Hyperscaler Frontier models will become less and less important for smart people that pursue a progressive adoption of local AI.
I think the future is probably local LLMs, but with some connection to superior models online for more complex requests. Maybe even renting computing power to them.
The problem is deciding which sub-task is easy without actually solving it.
I've been benchmarking models against 4,000 very subtle multiple choice science questions. gemma-4-26b-a4b-it is effectively as good as any other model I've tested so far (within a few percentage points), including Opus and Gemini Pro.
a router that can intercept the prompt in chat? as a token saving process
so how about how about xiaomi mimo and it's licensing?
all depends on training.
i keep waiting for the day a full state of the art model can be run in 1gb of ram 😅