I Burned Five Hours of AI Budget in One Session. Here’s the System I Built Next.
One evening I had a cloud model write me about 1,900 lines of TypeScript in a single session. It worked, and it also burned through my five-hour usage limit in a single sitting.
The limit wasn’t the real problem. The real problem was that the next session didn’t remember any of it. Why that file looked the way it did, which decisions were already made, what I’d told it never to do again — all gone. I got to explain it all a second time.
So I gave it a memory. This is what that looks like.
The problem: sessions start from zero
Claude Code loses its context when a session ends. Open it again and you’re strangers.
For one small project that’s fine, and I worked that way happily for a while. But once you have several projects, each with its own rules, and a growing pile of “never do this again” — re-explaining becomes the work. And when you forget to re-explain something, things break.
Mine broke. A cleanup script in a test run emptied a data file I’d spent an hour filling in by hand. There was no backup, and it wasn’t in git either because I’d gitignored that file myself. That was an hour I simply had to spend again.
The rule I wrote that night is still in my global file.
Three layers
What I settled on:
| Layer | File | Holds | Lifetime |
|---|---|---|---|
| Global | ~/.claude/CLAUDE.md |
Rules that apply everywhere | Permanent |
| Project | <project>/CLAUDE.md |
Rules for that folder only | Project lifetime |
| Memory | Individual .md files |
One fact per file | Permanent, edited individually |
Why split it: one big global file eats context on every single session, and project rules start contradicting each other. Layers mean only what’s relevant gets loaded.
What goes in the global file
Only things that hold no matter what I’m working on. In my case:
- Decisions that cost money. No paid APIs. Even when one is technically the better answer, I want the local option proposed first.
- Destructive actions. Never overwrite data I accumulated by hand — the rule that hour-long mistake bought me.
- Which machine does what. Some work belongs on the desktop, some on the always-on box.
- How to work. Plan first. Don’t re-ask about things already approved.
Things that change often — project lists, port assignments — don’t go here. That’s the next layer down.
Memory files: one fact per file
I have 17 of them right now. The rule is simple: one file, one fact.
---
name: short-slug
description: one line — this is what gets scanned for relevance later
type: user | feedback | project | reference
---
The fact itself.
**Why:** what made this a rule (the incident, and what it cost)
**How to apply:** what to actually do about it
The Why and How to apply lines are the whole point. A bare rule falls apart the first time an edge case shows up. Knowing why it exists means it can be applied to a situation I never wrote down.
The current split: 11 project, 3 feedback, 1 user, 1 reference. Most are just the state of ongoing work. The three feedback files are worth more than the rest combined — those are the times I said “don’t do it that way, do it this way,” and they stuck.
Hooks: what files can’t do
Files are passive, in that they only get read when something goes looking for them. Hooks interrupt on their own.
Mine run on every message I send, and they do three things:
- Inject the current time. The model has no idea what time it is. It needs that to work out what “tomorrow morning” means.
- Catch certain phrases and file them elsewhere. If I type something like “after work I need to check X,” it lands in a to-do file automatically.
- Drain queued messages. Notifications parked by other tools get folded into the session.
The point is that I don’t have to remember to mention any of it. Hooks don’t forget.
What actually changed
The biggest shift was where code gets written.
The rule I wrote after burning those five hours is one line: code generation goes to a local model, not the cloud one. The cloud model plans, reviews and integrates the result, but it doesn’t do the typing.
Because that rule lives in the global file, every session follows it without being told. I don’t say “use the local model for this” anymore, because I stopped having to.
What hardware that local model runs on is a separate post. Short version: one GPU I bought for gaming.
Don’t do these (I did)
- Don’t pile everything into the global file. Mine hit 63K and every session paid for it. Moving detail into separate reference files and leaving only pointers in the global one brought it to 24K.
- Don’t put code structure in memory. That’s what reading the code is for. Memory is for what isn’t in the code — why a choice was made, what must never happen.
- Don’t write “best practices.” A rule with no incident and no cost attached gets ignored later. Write down what it broke and what that cost.
- Don’t make hooks heavy. They run on every message. Slow hook, slow everything.
Still unsolved
- Old memories can go stale. The file stays exactly where it was while the fact underneath it quietly changes. I have no automatic way to catch that, so I delete them by hand when I notice.
- 17 files is manageable. I don’t know what happens at 100.
- This is Claude Code specific. Other tools structure this differently.
Closing
The bottleneck in AI-assisted coding turned out not to be the model. It was context — the same explanations, over and over.
A few files and a few lines of hook removed most of it. It’s less a system than a small stack of notes that don’t forget.