I Tried to Replace Claude With a Local AI Coding Agent. Here's Exactly Where It Broke.
A full day testing whether a local model running on consumer hardware can execute Claude-authored coding plans and cut token costs. It can, for one narrow kind of task. Here's the hardware math, the two real bugs I hit, and the actual capability ceiling.
I run a lot of side projects. Job boards, compliance dashboards, a writing app, a hurricane alert concept for my part of Mexico, a handful of others. All vanilla PHP 8, MySQL, and Tailwind, no framework, no ORM. I use Claude as my default AI pair programmer (Sonnet for day to day work, Opus for review), and it's cheap at my usage level, something like $7 a month on the API. But I kept seeing the same pitch everywhere: plan with a strong model, hand the actual typing off to a free local model, and cut your AI costs to near zero.
So I spent a day actually testing it, on my own hardware, on my own code. Here's exactly what happened, including the two places it broke and why.
The hardware
- Intel Core i9-13980HX (24-core)
- NVIDIA GeForce RTX 4070, 8GB GDDR6 VRAM
- 64GB DDR5 RAM
- 1TB NVMe SSD
- Fedora Workstation (I retired Windows 11 and Pop!_OS earlier this year; this is my only machine now)
The 8GB of VRAM turned out to be the single most important number in this whole experiment. Everything downstream, model choice, speed, even which bugs I hit, traces back to that.
Why local at all
The obvious first question: why not just point a local agent at my existing Claude subscription and skip local models entirely? Because that's explicitly against Anthropic's terms. In February 2026 Anthropic banned using Claude Free, Pro, or Max subscription OAuth tokens in any product other than Claude.ai and Claude Code itself. People were doing this to route their subscription through third-party tools, and Anthropic has been banning accounts over it. So the only two legitimate paths are the Claude API (pay per token, no subscription workaround) or genuinely local models. I went local, since that was the point of the experiment.
The agent: Hermes Agent + Ollama
I used Hermes Agent, an open-source coding agent from Nous Research that works like Claude Code but can point at any model, cloud or local. For the local model server I used Ollama, running as a systemd service on Fedora.
One thing worth knowing if you're on Linux: Ollama has no GUI on Linux at all, no app icon, no
window. It's a background service plus a CLI. The videos showing an "Ollama app" with a window are
running on Mac or Windows, where Ollama does ship a small native app. On Linux you just get
ollama run <model> in a terminal.
Picking a model: the VRAM math
With 8GB of VRAM, anything larger has to spill into system RAM over a much slower bus than the GPU's own memory. That's not a small speed penalty, it's the difference between fast and uncomfortably slow. So the naive plan was: pick the biggest coder model that fits entirely in 8GB.
That plan didn't survive contact with Hermes. Hermes enforces a hard minimum context window of 64,000 tokens to even connect a model, it refuses anything smaller outright with:
Error: context window below minimum 64,000 tokens
Most of the well-known small coder models, qwen2.5-coder:7b included, only support 32K context.
They fit the VRAM budget perfectly and still get rejected. That single requirement eliminated
almost every model that would have run fast on my card.
The models that clear both bars (fit reasonably on 8GB, and support 64K+ context) ended up being:
qwen3-coder:30b-a3b-q4_K_M, a mixture-of-experts model (only ~3B parameters active per token despite being called "30B"), 256K context, about 19GB on disk. Doesn't fit in 8GB VRAM, spills into the 64GB of system RAM.devstral:24b, Mistral's model built specifically for agentic tool-calling, 128K context, 14GB. Same story, spills into RAM.gemma4:12b, Google's general-purpose model, 256K context, 7.6GB. This one actually fits entirely in 8GB VRAM. Not code-specialized, but a solid pick for anything that isn't coding.
I also set the context length explicitly to 65536 tokens in Hermes's config rather than letting
it auto-detect. This matters more than it sounds: Ollama's actual default serving context isn't
the model's advertised max, it's based on your VRAM (about 4,000 tokens for anything under 24GB of
VRAM). If you don't override it, Ollama doesn't error when it runs out of room, it silently drops
the oldest messages, system prompt included, with no warning. That's a recipe for an agent that
quietly goes off the rails mid-task.
First real test: it looked like it worked. It didn't.
With qwen3-coder:30b-a3b-q4_K_M wired in, I gave Hermes a simple task:
create a test.php file that connects to MySQL and prints the connection status
Hermes displayed what looked like a tool call, formatted like this:
<function=write_file>
<parameter=path>
test.php
</parameter>
<parameter=content>
<?php
$servername = "localhost";
...
</parameter>
</function>
</tool_call>
It looked like a completed file write. It wasn't. Nothing was actually created. This turned out to be a real, currently open bug in Ollama itself: issue #17276, where Ollama's parser for Qwen3-Coder's tool call format intermittently fails on malformed XML the model emits, and instead of executing the call, it just leaks the raw text through as if it were a normal reply. No error, no crash, just a quiet no-op dressed up as a completed action.
I confirmed this by asking it to run ls -la test.php afterward. Same thing happened, another raw
<function=process> block printed as text, no actual command executed, no file on disk.
The fix: switching models
I swapped the wired-in model to devstral:24b, same Ollama endpoint, same 65536 context setting.
Reran the identical prompt. This time Hermes actually showed a real diff:
a/test.php → b/test.php
@@ -0,0 +1,18 @@
+<?php
+$servername = "localhost";
+$username = "root";
+$password = "";
+
+$conn = mysqli_connect($servername, $username, $password);
+if (!$conn) {
+ die("Connection failed: " . mysqli_connect_error());
+} else {
+ echo "Connected successfully";
+}
+mysqli_close($conn);
+?>
And the file actually existed on disk afterward, confirmed in the file manager. Tool calling worked for real this time. Devstral doesn't share Qwen3-Coder's specific parser path in Ollama, so it sidesteps that bug entirely.
The real ceiling: it's not a bug, it's the model
Feeling good about the setup, I asked Hermes (still on Devstral) to install a tool called CodeGraph, a legitimate, popular (71.8K GitHub stars) local code-indexing tool that plugs into agents like Hermes and Claude Code to reduce the number of tool calls needed to understand a codebase.
I gave it the GitHub URL in my first message. It asked me for the URL. I gave it again. It asked me to confirm it should clone the repo. I said yes. It replied:
I'm happy to help! How can I assist you today?
And dropped the entire task. No clone happened. This is not the same kind of failure as the Ollama bug. Tool calling was working fine at this point, technically. This was Devstral, a 24B parameter model, simply losing track of a conversation it was actively in the middle of, after receiving information twice and an explicit confirmation.
That's the actual ceiling of running a model this size locally right now: it can execute a single, precisely specified task well. It struggles to hold a thread across several turns, the exact skill that agentic coding work depends on most.
Would better hardware fix it?
My first instinct was: I need a Mac Studio with unified memory. Unified memory would genuinely fix the speed half of the problem, a Mac's RAM acts as VRAM, so a 30B+ model runs at consistent speed instead of splitting across a fast GPU and a slow bus into system RAM.
It would not fix the reliability half. The conversation-tracking failure I hit isn't a memory bandwidth problem, it's a model-scale and training problem. People running much larger models (70B-120B) on maxed-out Mac Studios report the same kind of thread-losing behavior compared to Claude, because none of the models you can realistically self-host right now, on any consumer hardware, are trained with the scale or agentic-specific reinforcement learning that Claude has behind it. A Mac Studio capable of running those larger models costs somewhere in the $3,000 to $6,000+ range. That's a lot of money to spend chasing "noticeably better but still behind," when the actual API cost of just using Claude directly is a few dollars a month at my usage level.
The verdict
Local models on consumer hardware are genuinely useful for one narrow thing: a single, well-specified task with no ambiguity, write this exact file, following this exact logic. They are not, right now, a drop-in replacement for Claude on anything that requires the model to track context across multiple steps, make a judgment call, or recover from a wrinkle mid-task.
If you're evaluating this as a cost-saving move, my honest advice is to test it on your actual multi-step work before committing to it, the way I did here. The savings are real for the narrow case. They're not a substitute for Claude on the work that actually takes the most time.
Appendix: exact setup reference
Models used and why:
| Model | Size | Context | Fits 8GB VRAM? | Verdict |
|---|---|---|---|---|
qwen2.5-coder:7b |
4.7GB | 32K | Yes | Rejected by Hermes (below 64K floor) |
qwen3-coder:30b-a3b-q4_K_M |
19GB | 256K | No (RAM offload) | Tool calls silently failed (Ollama bug #17276) |
devstral:24b |
14GB | 128K | No (RAM offload) | Worked correctly, hit model-capability ceiling instead |
gemma4:12b |
7.6GB | 256K | Yes | Reserved for non-coding/general tasks, not deeply tested |
Hermes provider configuration:
- Provider type: Custom OpenAI-compatible endpoint
- Base URL:
http://127.0.0.1:11434/v1 - API key: none required (blank)
- API compatibility mode: Auto-detect (resolved to Chat Completions)
- Context length:
65536(set explicitly, not left on auto-detect) - Reasoning effort: Skipped (neither model supports extended reasoning mode)
Tools enabled in Hermes: File Operations (read/write/patch/search), Terminal & Processes, Web Search & Scraping. Left disabled: plugins, additional MCP servers, messaging platform integration, full skill catalog seeding (kept at zero preloaded skills to protect the smaller context budget).
Key sources: