Timeline: rule hardened to no exceptions on 2026-06-26, first exception conceded 2026-07-01 five days later, rule rewritten 2026-07-15 after the GPU was tied up

I Wrote “No Exceptions” Into My Coding Rule. One GPU Broke It in Five Days.

My development rules contain one line:

Code generation goes to the local model. The cloud model plans and reviews.

On 26 June 2026 I added two words to it: no exceptions. That phrase survived five days. On 1 July I conceded the first exception and rewrote the line myself.

The rule itself came out of a night when a cloud model wrote 1,900 lines of TypeScript in one session and burned my entire five-hour usage limit in that sitting. I wrote that night up in the memory system post, so I won’t repeat it here.

There’s no metered API anywhere in my development stack right now. What gets billed is one subscription and one domain. This post isn’t about how good that feels. It’s a record of one thing: with a single GPU, the card decides how long a cost policy lives. And it’s about what’s printed on the back of a bill that reads zero.

Timeline: rule hardened to no exceptions on 2026-06-26, first exception conceded 2026-07-01 five days later, rule rewritten 2026-07-15 after the GPU was tied up

The stack, as it actually runs

Everything that’s running, and where it runs.

Job Where Meter
Planning, design, breaking work down Cloud model (subscription) Subscription limit
Writing and editing code Local Qwen3-Coder-30B-A3B (main PC, RX 9070 XT 16GB) None
Review, integration, build checks Cloud model (subscription) Subscription limit
Embeddings (semantic search) Local Ollama on a Mac mini, bge-m3 None
Image and video generation The same main PC GPU (ComfyUI) None
Hosting for this blog Self-hosted (Docker) None — the domain is the only paid part

The measured configuration is this:

  • Model: Qwen3-Coder-30B-A3B-Instruct UD-Q4_K_XL, 16.45GB
  • Server: llama.cpp llama-server, Vulkan backend, build b9587
  • Flags: -ngl 99 --n-cpu-moe 14 -c 16384 --parallel 1 --jinja
  • Speed: 41–50 tok/s generation, 180 tok/s prompt (pp256). VRAM 15.08GB of 15.9GB
  • Added 19 August 2026. I re-measured that generation figure on the same card with the same flags. It holds — but only when nothing else is using the GPU. With an idle image app holding a couple of gigabytes of VRAM, the same setup generated at 28 tok/s instead, and threw no error while doing it. The two conditions are measured here. What reproduced was the bottom of the 41–50 range; I never saw 50.

  • Editor: VS Code and Continue.dev on the Mac mini, pointed at the OpenAI-compatible endpoint on the main PC
  • Machines: main PC with 32GB RAM and an RX 9070 XT 16GB / Mac mini M4 16GB

Where those numbers came from is in the 9070 XT write-up, and how I landed on --n-cpu-moe 14 is in the tuning guide. I’m not explaining either one again here.

I measured one more thing. The model lives on the main PC and the editor lives on the Mac mini. That’s not localhost. Every request crosses a private mesh VPN on the way. End to end across machines I measured 45.9 tok/s, against 47 tok/s measured on the machine itself. For this workload the network hop was free. The firewall only accepts what arrives from the mesh.

Why the line sits exactly there

The rule says code generation only rather than “use local wherever possible,” and the reason is laid out as a role-split table in the 9070 XT post. In one sentence: the stage that burns the most tokens and the stage a local 30B handles well are the same cell. Move that one cell and the bill disappears. Force the rest across and the results get worse.

Added 18 August 2026. I later built an autonomous loop that follows this rule and logged every token in it. Orchestration turned out to be the bigger half. The driver’s output was 3.34× the code output, and 85% of the bill was cache. The measurements are here. The sentence above came from an interactive session, and it holds there. It didn’t hold for the loop.

I also don’t say this out loud every time. It sits in my global rule file, so every session follows it without being told, and I wrote up how that file is structured in the memory system post. What a session follows is the wording currently in that file. That wording changes once, further down this post.

The local model can’t read that file. So I keep the same guardrails a second time in the editor-side config — no paid APIs, never destroy accumulated user data, git rules, infrastructure protection. Maintaining one rule by hand in two places is a cost that ships with this shape. Change the rule and I change both. I haven’t automated that yet.

The procedure itself folds into one script. Opening the folder in VS Code turns coding mode on — it wakes the main PC, starts the server, and waits until the endpoint answers. After 20 idle minutes a scheduled task takes the server down and reclaims the VRAM.

What it actually cost

GPU (RX 9070 XT 16GB) Extra spend for this purpose: zero — a card I already bought for gaming
32GB RAM, CPU, Mac mini Extra spend: zero — machines I already had
Model weights 0 — a 16.45GB download
llama.cpp, Continue.dev, Ollama 0 — all open source
Setup 2026-06-07 first checks → 06-10 server up → 06-11 in operation (I never timed the work itself)
Cloud subscription One subscription. The amount isn’t in this post
Electricity I don’t know. I never measured it
API calls 0

Two of those rows stay empty on purpose.

  • The subscription price. I’m leaving the figure out, and I’m not dropping an estimate in its place.
  • Electricity. I’ve never put a meter on the wall socket. That’s why no sentence in this post says “X per month in power.” Local burns electricity, and how much is something I don’t know. When I measure it, I’ll write the number.

One thing is certain. The main PC is normally powered off. I wake it when I need it. Nothing here runs 24/7.

Where local loses — four places

1. Thirty seconds to the first token

The main PC is off. Starting a coding session means waking it over Wake-on-LAN, and the boot takes 28–30 seconds. Then the server loads 16.45GB, and during that the first request comes back 503. I never timed that load. What I know stops at “it recovers shortly.”

The cloud model waits zero. I hid this behind one script rather than removing it. A single command handles wake, start and wait, and 30 seconds is still 30 seconds.

2. One GPU, so one thing at a time

This was the expensive one.

The main PC has a single GPU. Gaming, coding and image or video generation all share that one card. So I put a lock file in front of it. The state is one of four — free / gaming / coding / imaging.

  • A watcher polls every 10 seconds, catches a game starting, and unloads the model.
  • If the lock reads gaming or imaging, the coding server refuses to start.

All of that works as designed. What came next is the problem.

I hardened the rule to “no exceptions” on 26 June. On 1 July I conceded an exception and rewrote the rule. The reason for that first exception wasn’t hardware — that day I admitted multi-file refactors and urgent fixes go better when the cloud model does them directly. Five days after I nailed it down.

Hardware came afterwards. A workload then had the GPU tied up on image generation, and through that stretch the coding model couldn’t start at all. Code that by the rule belonged to the local model kept getting written by the cloud one. An exception I allowed once became the default for that whole period.

On 15 July I rewrote the wording again. “Code always prefers local; fast iteration goes to the cloud directly.” I didn’t keep the rule. I edited it.

With one GPU, a principle loses to hardware. Anyone with two cards never meets that sentence.

I’ll answer one thing up front. I do own a second card, a 2080 Ti — the one I benchmarked in the two-card comparison. It sits in a different machine, it has 11GB, and generation on the same model was slower there. Moving the coding server onto it to dodge the contention is something I never tried. So the card this post counts is the one in the main PC, and whether a second card solves this, I don’t know.

3. A 16k context ceiling

Why the context is pinned at 16k, with 0.8GB of VRAM left over, is written up in the 9070 XT post. What belongs here is what that ceiling does to routing.

Work that needs a whole repository read can’t go to the local model. That stays with the subscription. So this stack isn’t “replace it all with local” — it uses both, by design, from the first day.

4. Handoff loses on multi-file refactors

The unit I hand the local model is one file, one change. Asking it to regenerate a full file in one go times out, so I chop the work smaller.

I tried a UI refactor spanning several tangled files this way and gave up. Running a search-and-replace handoff once per file was slow, and I didn’t trust the intermediate state where half the files had changed. That’s why the policy is mixed today. And when I do the typing myself, I write down that I did it myself.

Don’t do these (all of them mine)

  • Before writing “no exceptions,” look at what hardware the rule assumes. Mine stood on an assumption that one GPU was sitting free for coding. That assumption produced its first exception in five days, and once the GPU was tied up on image generation I rewrote the wording itself.
  • Don’t assume project rules reach the local model. It can’t read the cloud side’s rule file. I maintain the same guardrails by hand in two places.
  • Don’t call the server dead on one 503. It’s usually still loading the model. I wrote that down as a rule after confusing the two a few times.
  • Don’t push a large refactor through file-by-file handoff. I did, gave up, and went back to editing directly. When keeping a principle makes the result worse, the principle is the thing that should change.

Honest limits — what I never measured

  • I don’t know the electricity cost. I never measured watts at the socket. So this post reaches no conclusion that local is cheaper than cloud. What I know stops at there’s no metered API here.
  • I never worked out what the API bill would have been. I’ve never counted the tokens. I’d rather leave that number out than invent it.
  • I never measured how much less of the subscription limit gets used. Burning a five-hour limit in one session happened. So did the fact that it hasn’t happened since. The ratio between those two, I never measured.
  • One local coding model, and that’s all. I didn’t compare it against another one.
  • Quality still sits below frontier models. Where it splits is written up in the 9070 XT post. That’s why I’ve never deleted the review stage.
  • I never timed the model load. The 28–30 second boot is measured. What happens after that, I know only by feel.

So who should build this

  • Already own a GPU that sits idle? It’s worth doing. My extra spend was zero, and a 30B-class MoE genuinely runs on 16GB VRAM plus 32GB RAM. When VRAM is tight, n-cpu-moe tuning is the biggest lever available.
  • Buying a GPU in order to do this? Do the arithmetic first. How many months of subscription the card costs, and whether a 30-second wait and a 16k ceiling are livable for that many months. My card was already here, so I never had to run that sum.
  • If you are buying, VRAM capacity comes before brand. 16GB against 11GB mattered more than seven years of architecture — the table is in the two-card comparison.
  • One machine doing gaming and image generation too? Design the exclusion first. I bolted the lock on late, and only then did I see that my coding rule was standing on “the GPU is idle.”

Closing

There’s no API bill in my development stack. What’s there instead: a 30-second boot, a 16k context, one GPU that does one thing at a time, and a copy of a single rule maintained by hand in two files.

Zero doesn’t mean there is no cost. It means the cost arrives in a shape that isn’t a bill. I make that trade knowingly now. I didn’t know it before the night of 15 June, when I burned the limit.

Similar Posts