Can Local LLMs Replace Subscriptions?

October 10, 2026

Approximately 17 minute read time.

Happy Spoopy Season all!

For about a year or so, I had been subscribed to the $20/month plan for either OpenAI or Anthropic, depending on the month. There were so many cool things I was able to do with those subscriptions, like building mobile apps with little to no experience with mobile development, learning new skills rapidly, and playing around with Deep Research to see if an idea I had was worth pursuing, or if others have tried and failed at ideas that were similar.

And don’t get me wrong: OpenAI and Anthropic models are very good: Sonnet is typically my go-to model for most software engineering tasks, and OpenAI’s models have been really great at generating photos and helping me with other non-engineering related work like creating reports for me using MCP servers.

OpenAI and Anthropic aren’t the only players in the LLM game though: there is a menagerie of open-weight models in the market that, with the right local hardware, could easily replace the proprietary models I’m used to using. When looking at benchmarks that compare some of those models, like Qwen 3.8 or Muse Glimmer, to the proprietary frontier models, it’s quite impressive what some of those models are able to achieve.

Do they stack up to Opus, Sol, Fable, or Astra levels in terms of quality or accuracy? Probably not, but have I ever been able to use any of those models for a long enough time to actually make anything? Not on a $20/month subscription, no. $20/month would normally get me a solid hour of using Opus, at max, and then I’d have to wait until my limits got reset to continue. So me personally, I’ve pretty much gotten used to using the likes of Sonnet and Terra for most things.

By the way, if your job gives you some kind of AI tooling, do you know how much your usage is costing your organization? When I was working at Storable, I remember we were using Anthropic models via AWS Bedrock as opposed to subscriptions. However, in the instructions for the setup, they had originally recommended setting your default up to be Opus… Needless to say, we had to roll that waaaaay back.

Food for thought in case you really think you need Opus or Fable for everything. I’m here to tell you that you don’t.

Anyway, enough about subscriptions, let’s talk about local AI once again on this blog.

Making Local AI Possible

So if you’re a habitual reader of my blog, first of all thanks 😁😁😁, and you might remember how in My Primer with Hermes and in previous blog posts, I’ve been using so many different tools, and trying out a lot of different models. I’ve also spoken about how running local models has, in recent months/years, been difficult to achieve. And not only have I talked on this blog about that, but I also have talked to my wife a lot about different computers that I thought would be cool, and different deals that were around there…

Well, when I told her about a deal for the Asus ROG Z13 Flow that Best Buy had, I think she finally got tired of me talking her ear off about computers and LLMs and stuff and she pulled the trigger.

Bless her heart 🥲

But anyways, I am now the proud owner of a Z13 Flow with an AMD Ryzen AI Max+ 395 chip, and 64 GB of unified memory, and let me tell you something: as nice as this device is, that is literally the worst name you could come up with for any modern CPU.

Poorly named wafers aside, let’s discuss why this model in particular. Apart from the pricing being fantastic when this device was purchased, 64 GB of unified memory has turned out to be a really great sweet spot for running LLMs locally. This device is also perfect because of it’s portability: a desktop PC would have me tied to a desk whenever I wanted to use it.

But why this one and not a 128 GB model? Because they do have those, and that’s what every mini PC marketed for running local models has in it nowadays minimum. Honestly, I feel pretty comfortable in saying that 64 GB in 2026 is the perfect balance of price and performance: for what I’m experimenting with locally, 32 GB would not have nearly been enough, and I’ll show some numbers to back that up here in a bit. 128 GB would have been great, and I might have been able to run more powerful models, but there are really great local models available in the ~30B parameter range that I think a lot of people would be more than happy with using. 64 GB gives me enough space to run what I want at a price point that doesn’t cause my back to spasm.

While this might not be the most repairable device, like I’ve spoken about previously, this is an unfortunate reality of computer hardware and Generative AI use-cases in 2026: No VRAM? No local LLMs. Unless you’re planning on running your models in system RAM, you’re going to have a very poor user experience. It makes your definition of possible look really bleak compared to what that $20/month can typically get you.

Speaking of user experience and possible:

How Good is Possible?

Recall, if you will, the Hermes blog post I made, where I had tried to use local LLMs for coding and absolutely failed. The models weren’t large enough to know how to use tools, context windows were tiny, possible meant tiny models with short-term memory. Possible with 16 GB of unified memory, meant that I wasn’t able to achieve anything.

Well folks, I’m here to tell you, that when you quadruple the amount of unified memory you have, that your experience does, in fact, get much better using all these tools!

But what tools, you ask, am I actually using? What models? And what have I actually done with all that stuff anyway?

What is The Sharpest Tool in the Shed?

Well let’s talk about the tools first, then I can tell you what I’ve actually done with them.

So I first thought to myself: “Hey self! You like Ollama a lot right? Just use that and use that for everything!”. And I did, and I tried using Claude Code as a harness using Ollama, and uh… I’m not convinced at this time that Claude Code is all that good to use for local LLMs… I tried using some models with it (admittedly I forget which ones), and I was really struggling with actually using them.

One thing that I also began realizing as well, was that Ollama’s model library was somewhat limited. At the time of writing, Ollama’s models page had roughly 230 models available for local usage… Have you been to HuggingFace and seen the over 3.1 MILLION models they have? I rest my case.

So I started hunting around, and two names kept coming up: LM Studio Bionic, and Unsloth.

(LM Studio Bionic)[https://lmstudio.ai/] was ok: I really didn’t mind the chat functionality of it, or how it can create entire projects within the agent window itself. However, I didn’t like that it pretty much always wanted you to put all your chats into organized project folders, and, if you can believe it, I also didn’t like how limited the model quantization options were for various models.

Quantization, for those who aren’t aware, is the process of quantizing the weights used by an LLM down to specific precisions of bits, in order to make the model smaller to download and to fit into VRAM. With quantized models, you lose precision, but you gain slightly faster processing and more resource efficiency. Typically, 4-bit quantizations (or Q4 quants, as they are referred to) are the sweet spot: good precision at a fraction of the VRAM and disk consumption.

My problem with LM Studio, however, is their models page for, say, Qwen3.8-27B, provides me with 3 options: Q4_K_M, Q6_K, and Q8_0. All three are great, but that gives me no option to experiment with any other options. Another example: Muse Glimmer, Meta’s latest 30B open weight model: comes in one quant in LM Studio and it’s the Q4_K_M quantization. Again, great option, but I want more because I know elsewhere I can find more. I’m trying to experiment, not try the 5-10 models that LM Studio recommends.

Another gripe that I had about LM Studio was the CLI. Desktop applications are not my favorite way to use agentic coding tools: I love my CLIs. LM Studio’s CLI doesn’t currently have easy ways to start a coding tool like Claude Code or OpenCode easily. I was a little shocked at that, because even tools like Ollama have CLIs that do that!

But then I discovered Unsloth, and it had pretty much all the functionality I was looking for: varied model selections at varying quanitzations, an easy-to-use desktop application that feels intuitive to use coming from ChatGPT or Claude, and best of all, I’m an unsloth start claude or unsloth start opencode away from starting an agent with one of the models I downloaded.

The ability to start up an agent harness this easily is a must for anyone who is wanting to use their existing tools: Claude and Codex harnesses will typically have specific instructions or files that exist in places that typically don’t work with other agent harnesses. Being able to transition your models without having to uproot your entire workflow or configuration is a huge win.

In addition to the CLI, Unsloth also felt very familiar with what I had come to expect from other GenAI apps: Deep Research mode, Web Search tools, MCPs, but with the added benefits of all my models running locally, not requiring an internet connection, and with a plethora of free, open weight models available from some of the largest players like ChatGPT, Google, NVIDIA, and Alibaba, all the way down to small-time fine-tuners, distillers, and obliterators.

For me: Unsloth has been the sharpest tool in my local LLM shed.

Models + Tools = Success?

So here’s a harsh reality of the local LLM ecosystem that I haven’t quite discussed yet that I think is important to understand: getting into a groove with local models and tools, in a manner in which you can be productive, comes with a lot of trial and error.

If you have subscriptions to ChatGPT, Claude, Cursor, or another provider that offers some kind of agent harness, you’re likely accustomed to logging into your tools on your machine, sending over a prompt, and your agents producing the work you wanted it to produce in the same day. It all just works. You pay your subscription fee, you log in, and Opus one-shots a todo app for you in less than 20 minutes. Simple, easy, fast.

With local LLMs however, it’s a series of trial and errors, trying different combinations of models and tools, and seeing what works best in what. Depending on the size of the model as well, and how much computing power you have, it may not be nearly as fast as a subscription that you might have.

Let me give you a more practical example: you know that post I made about using ChatGPT to edit a Pokemon Gold save? Well typically editing those files in Claude or ChatGPT was relatively fast. Ultimately most of those saves had problems with them, but within about half an hour I would usually have one.

Care to guess how long it took Qwen3.8-27B to edit the save file? Well… I wasn’t able to edit the file inside the Unsloth Desktop application, because apparently it doesn’t accept .sav files. Instead, I popped OpenCode up with Qwen3.8 to have it edit the file. It took a whopping 5 hours and 11 minutes to edit the save file. I still need to play the game to see if it worked, and I still have to finish my posts about trying to get Claude to edit the saves that I’ve been lazy about getting out.

Another example: I wanted to convert my blog from Gatsby to a different framework. I started by having Muse Glimmer do Deep Research on good alternatives I could convert the blog to. Based on the research it had done, I decided that Astro would be a good choice! In terms of the Deep Research, I think local models do as good of a job as you might expect Deep Research to work in Claude or ChatGPT, so for that local LLMs are more of a toss-up.

When I started the conversion, I had started by trying to see if this model called Qwythos, a Qwen-based model trained on Claude Fable traces. It’s one of the few models that I can run on this machine that natively support a 1M context window. Now… Interestingly enough, this model didn’t always perform very well in a non-Anthropic harness. In Unsloth, it didn’t always produce the right syntax for tool calls within the desktop app. But in Claude Code, it performed very well!

Well enough to convert my blog in one shot? Ha, I wish 🫠

I had tried starting out in Plan Mode like I normally would in Claude Code for larger refactors like this, and Qwythos was able to, at the very least, come up with the plan. Not bad. But when it came to all the coding effort it took to actually convert the blog, that was a little more tricky: context window compaction was not as smooth within Claude Code, and Claude Code sometimes timed out waiting for responses from Unsloth during compaction, or even when I switched to the slower Qwen3.8-27B and waiting for responses from it. This did not end up being the set-and-forget refactor that would have been very capable for the likes of Opus or Sol.

However, after a few days of switching between Claude and OpenCode, switching to Qwen3.8 to finish things up, and multiple sessions later, this blog was finally completely re-written and migrated to Astro! No more Gatsby security issues and dependency hell for me!

So it just goes to show: your selection of the models and tools really does matter. Some models work really great: Muse Glimmer and Qwen3.8 have been extremely capable models, and I’ve even enjoyed the occassional use of Ornith and gpt-oss models as well! In contrast, there are models that I thought I would be really great to use, like GLM-4.7, that I have yet to find a harness where that model will work at all. Am I doing something wrong? Probably, I’m very new to this, but it just goes to show that doing this all locally doesn’t mean it’s easy or fast. What it does mean though, is that I can still delegate tasks to an agent for free.

Going Bigger?

So I mentioned earlier that I would maybe have some numbers detailing where my hardware typically ends up sitting at and with which models. The factors are numerous here, and as the technology around LLMs gets more advanced, I’m going to be excited to see how far we can push our hardware! But for now, my setup and the numbers look a little bit like this:

So my machine is the Z13 Flow with 64 GB of unified memory, as I mentioned before. In Armoury Crate (the (British?) tool the Z13 comes with for lower-level device management), I have 32 GB allocated to the CPU, 32 GB allocated to the GPU, and approximately 16 GB of the 32 GB of CPU RAM can be shared with the GPU. Not quite as seemless as managing RAM on a Mac, but not horrible.

Now: that 32 GB of GPU RAM is surprisingly very flexible when it comes to running any of these LLMs locally. One of the nice things that Unsloth has (that Bionic doesn’t), is that it has a little RAM estimator that can give you insight into how much RAM the model is going to take up. Very convenient for when you may want to mutli-task and you want to ensure that the model will fit within the dedicated GPU RAM.

Where it gets slightly more complicated is when we start talking about models that don’t fit within 32 GB of RAM. Windows 11 takes, on startup, with most apps disabled from opening on startup, approximately 9 GB to 11 GB of RAM, leaving about 20 GB in headroom to do other things, providing the LLM itself doesn’t take up too much more room in the system.

If you are someone who has been interested in local AI development, and you aren’t sure if a model could fit either into VRAM or system RAM, here is what Unsloth tells me about the amount of RAM each of the following models can fit into, with the given parameter counts, quantizations, and context windows:

ModelSizeQuantContext WindowUnsloth Estimated RAM
Qwen 3.827BQ4_K_M262,14437.3 GB
NVIDIA-Nemotron-3-Nano-Omni30B-A3BQ4_K_M1,048,57633.16 GB
Muse Glimmer30BIQ3_M131,07218.75 GB
Qwythos9BQ8_01,048,57650.71 GB
gpt-oss20BMXFP4131,07214.93 GB
Ornith 1.535B-A3BQ4_K_M262,14429.84 GB

But what about larger models like GLM-5.3 and some of the flagship Deepseek models? Well some of those are potentially trillions of parameters large, roughly 30x the size of the larger models that I’ve listed here. Even at a Q1 quantization, which would basically lobotomize the LLM, likely none of these models are capable of running on this machine.

Ok but what about models smaller than 1T but bigger than 35B? Well… Two things about that: first and foremost, those sizes of models rarely exist. Once you exhaust the inventory of models that exist between the 9B and 35B range (which is most of them that you would want to run to do any kind of coding or productivity task) the models basically jump to 200B or larger.

Second: those 200B models would never fit in this machine as LLMs currently exist, no matter what format or quant they come in. Say for example DeepSeek-V4-Flash-0731. That is a 284B model. On disk, the model at Q1 quantization is 83 GB. Most of the models that fit on this machine are roughly 11-20 GB on disk, but you can tell between the context window and the models sizes on disk that they take up a substantial amount of space in RAM and VRAM. Unsloth even has a little icon that shows up whenever you’re looking at the quantizations that tells you the likelihood of a model fitting on the machine. It’s got Green for should fit, Yellow for might fit, and Red for highly unlikely to fit. I can’t find any models that are 200B+ parameters in size that aren’t Red 🫠

Conclusion

Overall, these local models are turning out to be incredibly capable and useful. Granted, there are some quirks getting this all running, and it’s definitely not as smooth of an experience as getting a subscription from some of the other players in the AI game. But if you get the right mixture of capable tools, right-sized models, and you have reasonable ambitions, you can delegate a lot of the things you want to do to local models.

If you read this blog and you had more questions for me, feel free to reach out to me on some of the socials at the bottom of the screen! Have yourselves a great rest of the weekend, and have fun LLM’ing!