The server was fine. SSH was fine. The firewall rule was fine. The block was on a screen nobody was looking at.
I was in the middle of a measurement when two remote servers went unreachable at the same time. The coding llama-server on port 8080, and the occupier server I’d started on port 8081 for that measurement. 8080 was the same server that had been answering normally a few minutes earlier.
curl waited more than eight seconds and ended with Connection timed out. Not a 500, not a 404. The TCP connection itself never happened.
Everything I checked after that came back “fine”. The server process was alive, the request it sent to itself returned 200, the network path was 1ms, SSH never once dropped through the whole thing, the firewall rule that opens that port was still sitting there, and the security event log held no block record at all.
What was blocking it was a dialog box sitting in the middle of the screen. A window with an allow button and a cancel button, the kind a person has to click before anything moves. While that window waited for an answer, all of that app’s networking was blocked. And nobody was looking at that screen.
This isn’t the post I sat down to write
In the last post I measured a VRAM cliff and closed it like this. The moment I move the occupier’s -ngl from 1 to 2 the generation speed drops 29%, and the VRAM that moved across that gap is 70MB, so I can’t explain the cliff. I wrote down what came next, too. Push -ngl to 3, 4 and 8 and see whether the ground under the cliff is flat or keeps falling.
That’s what I was doing. The measurement failed in the end — background factors contaminated it and the data is unusable, and I’ll write that up separately. What I ran into in the middle of that failure was a far cleaner incident, and this post is the record of it.
I ruled things out in order
To keep observation and inference apart, here’s only what I checked, in the order I checked it.
| # | What I checked | Observation | What that settles |
|---|---|---|---|
| 1 | curl localhost:8080 from that machine, to itself |
200 | The server process is alive |
| 2 | From that machine, to its own remote IP | 200 | The listening binding is fine too |
| 3 | Tailscale ping between the two machines | 1ms, direct (no relay) | The path itself isn’t dead |
| 4 | SSH (port 22) at the same time | Never dropped | That machine is reachable from outside |
| 5 | The inbound allow rule for the dead port | Present. Same shape as another rule that had been working for hours | The rule didn’t get deleted |
| 6 | Security event log | Routine signature update events only, 0 block events | No block was written to a log |
Row 4 is the line in that table that drives a person up the wall. SSH was alive. So I was sitting inside that machine, investigating the fact that the machine couldn’t be reached from outside. The “the network died” hypothesis was already wrong at that moment.
With all six ruled out, one candidate was left. Something that leaves no log and exists only on the screen. So I took a screenshot.
What was on the screen
An interactive Windows Security pop-up.
“Do you want to allow this app to access public and private networks? — llama-server”
That dialog comes up in Korean on this machine, and the line above is my translation of it. Two buttons, allow and cancel, in the middle of the screen. While that window waits for an answer, every networking function of that app is blocked. Nothing gets written to a log. What was holding it isn’t a rule — it’s a decision nobody had made yet.
I’m not putting the captured image in this post. The tool captures the whole screen exactly as it is, and that isn’t something to upload as-is. The sentence above is all of it that moves over here.
I don’t know why it came up right then
The timing lines up. Just before this I killed the coding server process and started it again, for reasons that had nothing to do with this incident. That server had been up for weeks, and in all that time I had never once seen this pop-up.
My read: a freshly started process gets asked again, even when it’s the same executable. That’s the explanation that fits the order of events best. And it’s a read, not a measurement.
What Windows actually uses to decide “is this a new app” — the executable hash, the path, the first time it listens — I never confirmed in any documentation. I also never repeated it on purpose: kill the same server, bring it back, see whether the pop-up returns. This is an n=1 incident.
A terminal can’t click that button
An SSH session is a different session from the desktop session that dialog lives on. So there was nothing the shell could do. The order went like this.
- I captured the screen. I used a remote screenshot tool I already had — it runs a script in the interactive session through the Windows Task Scheduler. There’s a reason it’s built that way. An earlier measurement already showed me that processes SSH starts directly die when the session drops.
- I cropped the dialog’s buttons out of the captured image and read off their pixel coordinates.
- I got it wrong twice here. (below)
- I synthesised a mouse move and a click at the OS API level and pressed Allow.
- Both ports came back 200 within seconds.
The third item gets its own paragraph. The first attempt missed because I multiplied the screenshot’s display scaling wrong. The second attempt was dumber. On a multi-monitor setup the origin of the screen coordinate system isn’t (0,0). This machine has another monitor to the left of the primary one, so the virtual screen’s X origin is -2560. Hand a pixel coordinate from inside the screenshot image straight over as a click coordinate and it lands 2560px off. I clicked a completely unrelated tab once. Only on the third attempt, after querying the virtual screen origin separately and adding the offset, did I hit the button.
What it did to me, left in
1. I suspected the firewall rule first and added a new rule, then deleted it. It turned out to be completely unrelated. Once I found the real cause I cleaned that rule up. The symptom was “the port won’t open”, so looking at rules first was the right order — but adding another rule after confirming a rule was already there means I didn’t believe row 5 of my own table.
2. This incident cut off the measurement I was actually running. After fixing it I restarted the coding server and picked the measurement back up, and that restart itself is what stops the later data from lining up with the earlier baseline. The numbers I took that day are unusable as a whole. That’s why this post has no speed table in it.
3. I got the coordinates wrong twice. Exactly as written above. Scaling once, a negative origin once.
4. I don’t know exactly how many minutes this incident lasted. I was walking the diagnostic steps in order and didn’t leave dense timestamps. It felt like about ten minutes. There’s no exact value.
What I didn’t confirm
- I couldn’t narrow down the condition that made the pop-up appear. Whether “the process restarted” is the only candidate, or whether something else changed on that machine that day, I don’t know.
- I didn’t reproduce it. I never went back and killed it and brought it up again on purpose. This post is the record of one event, not a spec for a behaviour.
- Whether the same thing happens on another Windows version or another security configuration, I don’t know. One machine, one set of settings.
- Whether the networking was blocked from the moment that window appeared, or from earlier, I don’t know. What I observed goes exactly this far: it was blocked while the window was up, and it cleared the instant I clicked.
Don’t do these
- Don’t move on to “the network is fine” because the logs are clean. This incident left not one line in a log. It can’t. A decision nobody has made yet doesn’t produce a block record.
- Don’t read SSH connecting as that machine being fine from outside. What was fine was port 22, and nothing else. I sat inside it and floundered for a while.
- Don’t add another rule after confirming the rule is already there. I did. It was unrelated, and all it bought me was one more thing to delete.
- Don’t use screenshot pixel coordinates as click coordinates on a multi-monitor setup. If the virtual screen origin is negative, the whole click lands off by that much. On this machine it was
-2560, and that’s this machine’s number, not anybody else’s. - Don’t pull a rule out of this that says “Windows asks again every time you restart a process”. I never confirmed what it judges that on, and I never reproduced it. This is one event.
- Don’t assume there’s nothing on the screen of a machine you run headless. This entire post is what happens when that assumption is wrong.
Closing
There’s a shape this site keeps running into. In the compressed cache post the logs were clean and the result was bad. In the monitor post I read the same value five times and the value was fake. In the last VRAM post the numbers reproduced and didn’t explain the result. This time it went one step further. The cause was in a place my observation tools don’t reach.
Running a server remotely rests on one quiet assumption. That nothing is sitting on that screen nobody watches, waiting for an answer. I had never once checked that assumption. It had been true for weeks.
So I fixed one line of my checklist. When a remote service goes dead and the logs, the firewall rules and the network path all come back fine, the next step is to look at the actual screen. This time that step was dead last, and the six steps in front of it returned nothing but “fine”.
Money spent: zero. All I used was a machine I already had and a screenshot tool I’d already built, and this site pays for no APIs. What it cost instead was a whole measurement session that day.