Navigate Select ESC Close

Is Local AI Coding Actually Good?

2026-07-18 Education
15.0k
525
88
Tech With Tim
Tech With Tim
2.0m subscribers

Unlock all features

FREE: Get instant access to 10 AI summaries, chats, or transcripts per day.

Description

Get started with MindsHub Cowork for free (it's open-source): https://mindshub.ai?utm_medium=influencers&utm_source=youtube&utm_campaign=26q3-tech-with-tim&utm_content=integration-1 Everyone online tells you local models are the future and you should run everything locally. I ran the tests on high-end hardware — an RTX 4090 with 24GB VRAM and an M5 Max with 64GB unified memory — and I'm going to give you an honest answer. Want to make real money with coding? I share high-signal insights on careers, monetization, and leverage in my free newsletter. Join here and get my guide How to Make Money With Coding instantly: https://techwithtim.net/newsletter 🚀 Tools I Use Get 10% off with code techwithtim Openclaw setup: https://www.hostinger.com/techwithtim VPS setup: https://www.hostinger.com/techwithtim10 Wispr Flow (Best AI Dictation): https://ref.wisprflow.ai/TechWithTim-jul26 ⏳ Timestamps ⏳ 00:00 | Overview 00:26 | My PC Specs 01:21 | Cost VS Hardware Tradeoff 01:51 | Analyzing the Models 03:21 | MindsHub Cowork 04:33 | Models & LMStudio Setup 07:32 | Running Models in VSCode 08:32 | IDE & Harness Compatibility 09:53 | Local Model Speed & Results 16:01 | Comparison to Claude PAID Models 18:20 | My Honest Analysis Hashtags #LocalAI #AICoding #LMStudio UAE Media License Number: 3635141

Top Comments (10)

@TechWithTim 2026-07-16

Get started with MindsHub Cowork for free (it's open-source): https://mindshub.ai?utm_medium=influencers&utm_source=youtube&utm_campaign=26q3-tech-with-tim&utm_content=integration-1

1 1 replies
@salat 2026-07-18

Drop the middle man: Learn to use llama.cpp directly - Router mode or with llama.-switch..

25
@cronocritcal6490 2026-07-18

The biggest thing on offer by local models is privacy. Using enterprise cloud models means sharing whatever data you are working on with a 3rd party which is some cases, isn't feasible due to laws, regulations, or company policies.

14 1 replies
@wooof8153 2026-07-18

On two DGX Sparks running Deepseek V4 Flash 284b full native precision (MXFP4 + FP8 QAT model), 60 tok/s decode, 2k tok/s prefill, and ~2.2M tokens of total kv-cache right now. It's been shockingly stable in different 24/7 agent and is very good at non-vision coding tasks. Can't wait for the final checkpoint releasing soon with vision.

16 5 replies
@Joseph-ur 2026-07-18

You are not limited to a model that fits in VRAM, llama.cpp + MoE model + MTP = sweet

13 3 replies
@Weißesnashorn 2026-07-18

Doesn't feel like you've answered the central question of the video. . You created a oneshot game and 1 of 2 didn't work an you say the reason might be the harness / model compability. So what should be the conclusion? My experience: Without promoting local AI, I make totally different experience 🙄 . That video is more a: Look, I have spent 10 minutes with a local model and this is what I have collected! Little bit of a timewaster...

9
@PhilippDer2Mac 2026-07-18

nice video but let me correct you on something. Several points, but I'll go through one at least One of them being the CPU/GPU divide or in other words loading a part of your model on RAM. It's totally doable and you won't be facing 2-3t/s at all (unless you use your swap for this). I run my context window and my initial run of qwen3.6 (24gb) in size with 150K context window on my 8GB VRAM and the rest is being loaded on my 32GB RAM and I hit more often than not around 40-50 t/s. With the context window full it may drop to 20 t/s but it's very well above what I expected. The tricky part is to not use ollama or lm studio, but llama cpp (with llama swap) and on linux (dual boot laptop)

10
@Snowsea-gs4wu 2026-07-21

11:51 like that joke: - quick: add 1+1 - 3 - what?!?! - did you want speed or precision? LOL😂

0
@michaelteegarden4116 2026-07-19

Thank you, sir! I appreciate honest eyes-open takes on running open source LLMs on the hardware setup quality that most of us actually have. "Get the A.I. model for the rig you actually have, not the rig you want to have." :)

0
@reynolrodriguez4982 2026-07-21

another important aspect is the license, just a few are really free to use for making profit

0

Unlock the Data Inside
Turn Videos into Knowledge

  • Get FREE 10/day: transcripts, summaries, chats
  • Chat with videos, export text & PDF
  • $1 free API credit for RAG, chatbots & research

Free forever plan • All features unlocked

App screenshot