On-Device Translation Turned “It Doesn’t OOM” Into “OOM Doesn’t Work”
I started writing this blog in English, which meant showing up on Reddit and Hacker News. Reading was fine. Writing was the bottleneck.
Every comment went through the same loop. Think it in Korean, convert it to English I’d be willing to post, check it again. For a two-paragraph reply that loop cost me minutes, and I was doing it for single comments.
So I looked at automating it. Before automating it, I measured whether I should.
The answer was no. Not because anything was slow. The fastest candidate returned grammatically perfect English in 0.41 seconds, and in six of ten cases it had changed what I meant. Nothing errored. Every request succeeded.

Two things wrong with this experiment, up front
I’m putting the flaws before the results, because both of them are the kind a reader should get to weigh themselves.
First: one of the contestants was the judge. One backend I tested was claude, running as a CLI subprocess. The thing that read the outputs and called the failures was the same tool. It scored a competition it had entered. There’s no way to dress that up.
Second: I never scored anything. The design had a four-axis rubric — naturalness, nuance retention, community fit, factual integrity — on a 1–5 scale, with claude as the baseline. The results file generates a scoring table per case. I opened it. Every cell is blank. It was built to be filled in by hand and nobody filled it in. So the pass mark I’d written down — “within 1 point of the baseline counts as usable” — was never once applied.
That’s why there are no scores and no rankings in this post. What’s left is a table of my Korean next to the English that came out. The failures below don’t need a judge. A meaning got reversed, an output broke, a fact changed. Anyone who reads Korean can check every row against the raw file.
That’s the part I trust, so that’s as far as this post argues.
What I measured
| Item | Value |
|---|---|
| Direction | Korean → English, one direction only |
| Cases | 10 |
| Backends | claude · apple (Apple Translation, on-device) · ollama:qwen2.5:7b |
| Hardware | One Mac mini, all three |
| Repeats | 1 |
| Outcome | 30 of 30 completed, zero errors |
The ten cases aren’t small talk. They’re all in the register I actually write in — technical community answers that share measurements honestly. I lifted them from the answer cards I’d already written for Reddit. These were sentences I intended to post.
Each case has a trap written down before the run. Case 10’s says: “the subject is entirely absent — English requires one, so the model is likely to fill in ‘you’ or ‘it’ wrongly.” That matters. Defining failure after seeing the output lets you fit the definition to whatever came back.
The traps all cluster in one place: the spots where Korean hedging, softening and dropped subjects break on the way into English. Phrases like “I haven’t measured it, but”, “14 isn’t the answer”, “it looks like it might be”.
Speed
| Backend | OK / fail | Median | Range |
|---|---|---|---|
apple |
10 / 0 | 0.41s | 0.33 – 2.27s |
ollama:qwen2.5:7b |
10 / 0 | 1.98s | 1.32 – 14.58s |
claude |
10 / 0 | 4.09s | 3.33 – 5.41s |
apple wins by an order of magnitude. It’s on-device, so there’s no network in the path at all.
If I’d stopped here I’d have picked it. Free, offline, 0.41 seconds, and a clean sheet — zero failures out of ten. Read the table on its own and there’s nothing to think about.
Only
ollamalogged time-to-first-token and generation rate (median TTFT 0.45s, 20.6 t/s). The other two don’t stream, so those fields are empty. Don’t compare the three on that axis. The table above is end-to-end wall clock and nothing else.
Then I read the output
One note on how to read these. The left column is my own English rendering of what I wrote in Korean — it’s there so the comparison is legible, and it’s not neutral evidence. The right column is verbatim from the results file. The two headline cases further down show the Korean itself.
apple — six
| What I wrote (Korean) | What came back | What broke |
|---|---|---|
| “It doesn’t OOM, but sometimes it suddenly slows down” | “OOM doesn’t work, but there are cases where it suddenly slows down” | Meaning reversed |
| “13 and below OOM’d” | “13 or oss is OOM” | Output broke |
| “I haven’t been able to run ROCm” | “I haven’t been able to play ROCm” | run → play |
| “I burned two days” | “it took two days” | Wasted → spent. The regret is gone |
| “it just picked it up” (was detected) | “I just got caught“ | Subject became me |
| “I haven’t measured it, but” | “I haven’t seen it again“ | Unrelated sentence |
ollama:qwen2.5:7b — four
| What I wrote (Korean) | What came back | What broke |
|---|---|---|
| “Network requests are zero, and the AI features are optional so they’re off by default” | “Network requests are off by default and AI features are optional” | Claim altered |
| “I’ve got no basis for saying Vulkan is faster” | “I haven’t tried ROCm, so I can’t say it’s faster than Vulkan” | Comparison flipped |
| “fuzzy search” | “locally spreads search” | Read as a verb |
| “I haven’t measured it, but” | “Re-running it wasn’t necessary“ | Didn’t → didn’t need to |
The line that scared me
Mine: OOM은 안 나는데 갑자기 느려지는 경우가 있어요.
(It doesn’t OOM, but sometimes it suddenly slows down.)Output: OOM doesn’t work, but there are cases where it suddenly slows down.
I was saying the out-of-memory error never fires. What came out says the OOM mechanism is broken.
Picture that posted to r/LocalLLaMA. The grammar is clean. There’s no typo. A reader puzzles over it for a second and scrolls on. And in that thread I’m now someone who doesn’t quite know what he’s talking about.
What makes it worse is where it landed. “No error fires, it just gets quietly slower” wasn’t a throwaway line — it was the whole point of the answer. Of ten sentences, the one it reversed was that one.
And the one where the claim changed
The ollama failures have a different shape. Nothing reads awkwardly. The claim is just different.
Mine: 네트워크 요청이 0이고, AI 기능은 옵션이라 기본은 꺼져 있습니다.
(Network requests are zero, and the AI features are optional so they’re off by default.)Output: Network requests are off by default and AI features are optional.
I said network requests are zero. That’s an absolute, and anyone can verify it.
What came out says they’re off by default — which means there’s a switch.
The extension I built doesn’t make network requests at all. That’s the entire product. The translation turned a structural guarantee into a preference you could flip. To someone choosing it for privacy, those are two different pieces of software.
In the same sentence, “fuzzy search” became “locally spreads search” — it read fuzzy as a verb. That one’s just funny. The first one isn’t.
None of them said they’d failed
Thirty of thirty completed. No exception, no timeout, no empty response. Read the log and this is a clean run with a 100% completion rate.
I’ve seen this shape before. In the fourteenth post an idle app held VRAM and my local model dropped to two thirds speed with no warning in the log. Every request came back 200.
Translation is worse than that. When something gets slow, you can at least feel it. When a translation is wrong, catching it means reading the output and checking it against the source.
And if I can do that, I didn’t need the translator.
A tool I can only use safely when I’m able to audit it isn’t a tool for me. It’s a tool for someone who already doesn’t need it.
What I didn’t measure
- I ran it once. No repeats. The same sentence could come back differently.
- Ten cases. That’s a set of examples, not a statistic. Don’t turn six and four into a failure rate.
- One direction. Reading — English into Korean — wants the opposite things: real time, full screen, good enough to follow. None of this transfers to that.
- I never tested ordinary sentences. I deliberately picked technical hedging. How
applehandles “where’s the bathroom”, I don’t know. Probably well. - One Mac mini. Other hardware, I don’t know.
- I didn’t tune anything.
ollama:qwen2.5:7bran stock. A better prompt might fix some of this. I didn’t try. - The judge was a contestant, and I never filled in the scores. Both as described at the top.
Don’t do these
- Don’t read this as “Apple Translation is bad.” Ten sentences, one direction, one run, and I picked the hard ones on purpose. That’s a worst-case portrait.
- Don’t pick from the speed table. I nearly did. On that table
appleis flawless — right up until you read what it produced. - Don’t treat a translator’s success count as a quality signal. Mine was 30 of 30.
- Don’t apply this to a language you can’t check yourself. I read both, which is the only reason I have a table. In a language I couldn’t read, all ten of those failures would have shipped.
What I did instead
I’m not automating the writing. I still do it by hand.
Four seconds was never the problem. I post a comment every few minutes at most, and 4 seconds versus 0.4 changes nothing about my evening. The problem is that I have no way to know when it’s wrong. The fast one is the one that breaks, it breaks silently, and catching it puts me back to reading both versions.
Here’s what the evening cost. I built the benchmark harness, wrote the ten-case set with its traps, and compiled and signed a command-line wrapper around Apple’s translation framework so it could be scripted at all. All of it to conclude I shouldn’t build the thing. I don’t regret it — that’s what measuring first is for — but it’s the honest price tag.
Reading is still open. The direction is different and the requirements are inverted, so it needs its own measurement. I haven’t run it.
One footnote on cost. The baseline, claude, is a session tool that already runs inside my subscription. This site doesn’t use paid APIs. The bill for this one was zero.