The compressed KV cache was already half speed at 7k, and died at 32k
A title went up on r/LocalLLaMA: “Qwen3.8 27B at 50 tok/s with 100k Context on a 16GB GPU!”
100k of context on 16GB, with a 27B model, at 47–50 tokens a second. Those three figures are the OP’s report, not my measurement. The method is KVarN, a feature of the llama.cpp fork beellama.cpp that tiers the KV cache by precision and is described as fitting up to 7.5× more cache in the same VRAM. The card in that post is an RTX 4070 Ti SUPER, 16GB.
My card is 16GB too. An AMD RX 9070 XT.
One claim, three numbers — 7.5×, 50 tok/s, 100k, all of them the OP’s. This site doesn’t quote a performance claim it has no repeated measurements for. So I installed it and ran it myself.
The result first. With 7,168 tokens sitting in the context, the compressed cache ran at half the speed of the fp16 cache. 9.04–9.31 tok/s became 5.27 tok/s. Then I pushed depth to 32,768 and the process died with exit code 0xC0000409. I never got near 100k.
The hole in this post, before any of the numbers
This sits right after the lead deliberately. It’s the biggest hole in the post, and it’s the thing to have in hand before reading a single figure below.
I checked the OP’s actual command line after all the measuring was done. It contains this.
--spec-type draft-mtp
That’s MTP speculative decoding. The model drafts its own tokens with its own MTP head and then verifies them, so no separate draft model is needed. It raises generation speed directly, and the OP’s 47–50 tok/s includes it.
My numbers don’t. The tool I used is llama-bench, which measures pure forward-pass throughput and never runs the accept/reject loop that speculative decoding lives in. Measuring that means standing up a server and attaching a generation loop, and I didn’t set that up this round.
So here’s the exact reading. I measured “kvarn compression plus flash-attn” on its own. The first reason my generation numbers sit below the OP’s isn’t that the claim failed to reproduce. It’s that I left MTP switched off.
One thing keeps this post standing. Every A/B/C comparison below happened inside my own card, with the same tool and the same flags. The difference between compression on and compression off has nothing to do with MTP. And the crash needs no interpretation at all.
The option name had already changed once
I got it wrong at the first step.
Names like TurboQuant and TCQ circulate in the Reddit thread and the project description. I went through the documentation for those names and found a redirect written into beellama-args.md: the legacy turbo2/turbo2_tcq now point at kvarn2. The option names had already turned over a generation.
Then I made my second mistake. I read the docs and picked kvarn2, the most aggressive setting. It has the highest compression ratio, so I guessed that had to be what reproducing the claim needed.
Wrong. Here’s what the OP actually used.
-ctk kvarn5 -ctv kvarn4
kvarn5 for K, kvarn4 for V. The two sides aren’t symmetric. By the OP’s own account both started at kvarn5 and reached 88k, and dropping V one step to kvarn4 saved 6% VRAM and filled out the 100k.
Opening the documentation and checking the command the original author actually typed turned out to be two different jobs. If I’d run with the kvarn2 I guessed at, I’d have published a “doesn’t reproduce” verdict from the wrong configuration. So I rebuilt this round around kvarn5/kvarn4.
I threw the entire first round away
I ran all of it on 31 August. Eight probe sweeps plus three main sets, every log generated, the script finishing on its own.
I opened it the next day and the jsonl files held no results at all. What sat in the stdout log was this.
usage: llama-bench.exe [options]
And one line in stderr.
error: invalid parameter for argument: -c
This build of llama-bench has no -c flag. That isn’t how context size gets passed to it. Every process printed its help text and exited cleanly. Nothing was wrong as far as the script could tell, and exit_code was recorded as null. Eleven pairs of log files, all created, every one of them containing --help.
I rebuilt it to set context with -d (n-depth) instead of -c and ran again on 1 September. One evening went into collecting help text. A clean log and a completed measurement are different things — I’ve already written that twice on this site, and it caught me again.
What I measured
| Item | Value |
|---|---|
| Hardware | Ryzen 7 5700X3D · AMD Radeon RX 9070 XT (16GB) · 32GB RAM · Windows 11 |
| Hardware in the original post | RTX 4070 Ti SUPER (16GB, NVIDIA) — a different GPU vendor |
| Backend | Vulkan (ggml-vulkan.dll), AMD proprietary driver, matrix cores: KHR_coopmat |
| Baseline engine | mainline llama.cpp, build 9587 (d2e22ed97) |
| Comparison engine | beellama.cpp v0.4.4 (released 2026-08-29), build 11573 (f8cd4e6dd), Windows Vulkan build |
| Model | Qwen3.8-27B-i1-IQ4_XS-GGUF-Smaller.gguf — the exact file the original post names, no substitution. 13.5GB, attention IQ4_XS / FFN IQ3_S hybrid |
| Compression flags | -ctk kvarn5 -ctv kvarn4, --kv-tail-tokens 1024 |
| Flash Attention | -fa on — identical across all three configurations (variable held fixed) |
| Chat template | none (see limits below) |
| Tool | llama-bench.exe, the copy shipped with each build |
| Repeats | 3 per configuration (-r 3) |
The three configurations are these.
- A — mainline llama.cpp, fp16 cache. The baseline.
- B — beellama.cpp, fp16 cache. Compression off. This one is the sanity check.
- C — beellama.cpp, kvarn5/kvarn4. The OP’s setting.
B is the heart of this experiment. If B differs from A, then I’m looking at the fork’s own overhead and I can’t attribute anything I see in C. Only if B matches A can C’s difference be pinned on the cache.
Results
Everything below is three repeats, reported as the median and the full range of the three. llama-bench prints its own average for each run too, and those land in the same place.
depth 0 — nothing in the context
| Configuration | pp tok/s (median / range) | tg tok/s (median / range) |
|---|---|---|
| A: mainline, fp16 | 101.29 / 101.20–101.42 | 10.88 / 10.86–10.88 |
| B: beellama, fp16 | 102.10 / 101.54–104.12 | 11.15 / 11.08–11.15 |
| C: beellama, kvarn5+kvarn4 | 106.31 / 106.23–107.56 | 11.66 / 11.61–11.72 |
A and B sit within 1–2% of each other. Sanity check passed. Moving to the fork on its own does nothing.
And C is the fastest of the three. The configuration with compression on came out ahead of the two with it off.
If I’d stopped here I’d have written “KVarN is free.” Read that table alone and that’s what it says.
But C’s jsonl carries these fields alongside the numbers.
kv_tail_tokens: 1024
kv_tail_tokens_effective: 128
The compression logic touched 128 tokens on that generation run. The cache is empty, so there’s nothing to compress. What I measured at depth 0 isn’t the cost of compression. It’s compression not doing anything yet.
depth 7168 — with the cache actually filled
| Configuration | pp tok/s (median / range) | tg tok/s (median / range) |
|---|---|---|
| A: mainline, fp16 | 100.88 / 100.62–102.18 | 9.04 / 9.02–9.07 |
| B: beellama, fp16 | 99.18 / 97.20–101.29 | 9.31 / 9.28–9.32 |
| C: beellama, kvarn5+kvarn4 | 60.33 / 58.89–60.63 | 5.27 / 5.25–5.28 |
At this point kv_tail_tokens_effective reads 1024. The tail is full and the compression path is running flat out.
Generation fell from 9.04–9.31 to 5.27. Prompt processing came down from around 100 to 60.33. A and B hold in the nines side by side at the same depth. This isn’t the engine. It’s the cache.
The range across the three repeats is 5.25–5.28. That’s a spread of 0.03. It isn’t a stray reading.
Read the numbers straight and they say this. The feature I turned on to fit more context cut my speed in half once the context got long. That points the opposite way from “7.5× more with almost no loss” — on this card, this backend, this quant.
Then I pushed depth further
No repeats here, one pass each. And the batch parameters differ from the tables above — those ran -p 512 -n 128, this one ran -p 16 -n 16.
| depth | pp tok/s (n=16, single run) | tg tok/s (n=16, single run) |
|---|---|---|
| 16384 | 11.53 | 4.40 |
| 32768 | crash | — |
Don’t subtract the numbers in this table from the ones above. Different batch sizes mean it isn’t the same axis. Two things are readable here and nothing else: the direction, and what happened at 32768.
At 32768 the process exited with this code.
exit code -1073740791 = 0xC0000409 = STATUS_STACK_BUFFER_OVERRUN
That isn’t running out of VRAM. There’s no OOM message anywhere. The stack was corrupted and Windows killed the process, and that exit code is sitting in vram_log.csv exactly as written.
Which depth it dies at, I don’t know. I never narrowed the gap between 16384 and 32768.
I failed to measure VRAM
This is the part of the run I regret most. I wanted to verify 7.5×. That’s one of the three claims.
I failed. Here is every artefact I have left.
| tag | exit | vram_before_MB | vram_after_MB |
|---|---|---|---|
| A_baseline_fp16 | 0 | 2738.7 | 2788.8 |
| B_beellama_fp16_sanity | 0 | 2788.8 | 2781.4 |
| C_beellama_kvarn54 | 0 | 2781.4 | 2785.4 |
| D_kvarn54_depth_scaling | -1073740791 | 2793.4 | 2788.3 |
All of it reads 2.7–2.8GB. These are runs that loaded a 13.5GB model and executed it. Those values are adapter memory before the process came up and after it finished — an idle desktop reading. Not one peak during a run got captured.
So this post has no VRAM table, and I say nothing about 7.5×. I didn’t disprove it and I didn’t confirm it. I didn’t measure it.
The method was crippled before that anyway. There’s no nvidia-smi on an AMD card, so I used the Windows performance counter \GPU Adapter Memory(*)\Dedicated Usage — an OS-generic counter, not a vendor tool. The per-process counter (\GPU Process Memory(*)) reported 32GB for a single process during pre-checks. On a 16GB card. That value is physically impossible, so I dropped that whole family. I didn’t use Win32_VideoController.AdapterRAM either — 32-bit overflow makes it report a 16GB card as 4GB. Total VRAM in this post comes off the product spec and nowhere else.
One line of arithmetic is all I’ll leave. From the model config (num_hidden_layers=64, num_key_value_heads=4, head_dim=256), an fp16 KV cache costs 262,144 bytes (256 KiB) per token. That’s a calculation, not a measurement. How the kvarn5/kvarn4 pair arrives at 7.5× isn’t spelled out in the documentation.
What I didn’t measure
Writing the limits out at length is how this site works. This time the list runs long.
- I left MTP speculative decoding off. The one from the top. It’s the biggest hole here, and on its own it wrecks any absolute comparison against the OP’s numbers.
- I never saw the Reddit original as a primary source. Access to
reddit.comandold.reddit.comwas blocked in this session’s tooling. I confirmed the OP’s command line from two secondary summaries indexed by search engines. Both give the same details independently (kvarn5/kvarn5 → kvarn5/kvarn4, 88k → 100k), which raised my confidence and did nothing else. No screenshot, no permalink I checked with my own eyes. - The GPU vendor differs. NVIDIA against AMD. Whether the KVarN kernels are tuned mainly for the CUDA path, and how far kernel maturity on the Vulkan path diverges by vendor, I don’t know. I didn’t read the source and I didn’t measure it. The numbers in this post are a result, not a cause.
- One card, one driver version. This result covers Adrenalin 32.0.31041.1004 and nothing beyond it.
- beellama.cpp v0.4.4 was four days old when I wrote this. It’s one release of a fast-moving fork.
- I applied no chat template.
llama-benchis a tool templates never enter. The original post’s point about cutting thinking tokens is about perceived response time, and what I measured is raw tok/s. Different axis. On top of that, whether the OP’s 50 tok/s is final-response or raw generation, I don’t know. -dis a simulation. It doesn’t process a real 7,168-token document from the start — it fills the KV cache to that depth and measures from there. Whether that matches real use, I didn’t check.- VRAM failed, as described above.
- I never narrowed the crash point. Somewhere between 16384 and 32768, and that’s all I have.
- The depth-scaling run is a single pass. With no repeats, 4.40 is an anecdote and not a measurement.
- I never confirmed other apps were fully closed. The GPU lock reading
freeand zero processes touching the GPU are two different statements. A browser was probably up.
Don’t do these
- Don’t read this as “AMD is slower than NVIDIA.” There’s no data for that here. The GPU generations differ, the vendors differ, and the sample is one card on each side. Pulling that sentence out of this post is manufacturing something that isn’t in it.
- Don’t read it as the OP inflating anything. I think that person saw those numbers on that card. On top of which, I measured with one of their features turned off. The point of this post is “I can’t put a number into my own plans without repeated data and reproduction conditions”, not “that number is fake”.
- Don’t conclude beellama.cpp or KVarN is bad. The sanity check, B, landed within 1–2% of mainline. The engine is fine. What I saw is how a compressed cache behaved on one GPU, through one backend, on one quant, in one release. Stretching it past that isn’t allowed.
- Don’t call TurboQuant/TCQ a name that never existed. It was redirected to
kvarn2as versions moved on, and the documentation says so. - Don’t choose from the depth 0 table. I nearly did. On that table the compressed cache is the fastest thing there, because
kv_tail_tokens_effectivewas 128.
What I did instead
It’s not in my working setup. I’m still running the fp16 cache.
Half the speed isn’t the reason. Half speed for genuinely ten times the context is a trade I’d take. The reason is that it dies at 32,768, and that it dies as a stack overrun rather than an OOM. When something dies for want of memory, I know how far I can go. When the stack gets corrupted, I don’t.
And this post leaves one question it couldn’t answer. What happens with MTP speculative decoding on. llama-bench can’t measure that. It needs a server and a generation loop attached, and I haven’t done it yet. Competing with the OP’s 50 tok/s for real means measuring that. Until then all I hold is a compression-on against compression-off comparison, and this post claims exactly that far and no further.
This is the same shape as the fifteenth post and the fourteenth. All three times the logs were clean. The translators went 30 for 30, the GPU came back with 200s, and this time a bench script left eleven pairs of logs and ran to completion. The absence of errors was worth nothing.
The spend was zero. The card was already in that machine, and both engines and the model file are free. This site doesn’t use paid APIs. What this one cost was two evenings, and one of them went into collecting help text.