Dot plot of top-1 cosine similarity for two embedding models with a 0.75 cutoff line

Both Embedding Models Were Right. Only One Survived My Threshold.

I sat down to confirm something I remembered. Months ago, when I added local semantic search to a bookmark extension, I was sure one embedding model had fallen apart on my non-English bookmarks. I swapped models, it got better, I moved on.

So I rebuilt the test to write it up. It didn’t reproduce. Both models handled everything I threw at them, in both languages, without a single miss. That cost me an evening of measurements and the post I had already outlined, which I threw away.

What I found instead is worse, because it doesn’t announce itself at all.

The setup

24 bookmark-style titles, split across two languages. 8 queries, also split across both. Two embedding models running locally through Ollama:

  • nomic-embed-text — 768 dimensions, about 300MB
  • bge-m3 — 1024 dimensions, about 1.2GB

For each query I embedded everything, took cosine similarity against all 24 documents, and looked at what came first.

Result 1: both models were right, every time

8 out of 8 for both. Every query returned the correct bookmark in first place, including the non-English ones on the smaller English-oriented model.

So my memory was wrong, and I’ll take that trade — I would much rather find out here than after publishing it as a finding somebody else acts on.

Result 2: the numbers live in different worlds

Then I looked at the actual similarity values.

Dot plot comparing top-1 cosine similarity across two embedding models for eight queries, with a 0.75 cutoff line: five of eight pass on nomic-embed-text, zero of eight on bge-m3
nomic-embed-text bge-m3
Correct result ranked first 8 / 8 8 / 8
Similarity of that result 0.697 – 0.911 0.588 – 0.747
Highest score seen 0.911 0.747
Same 24 documents, same 8 queries, same machine. Only the model changed.

Look at where those ranges sit. The best score bge-m3 ever produced (0.747) is lower than the worst score nomic produced (0.697)… barely. The two ranges almost don’t overlap.

The ranking was identical in both cases, and the numbers sitting behind that ranking were not even close.

The trap

Here’s the line of code that gets written in every semantic search project:

if (similarity > 0.75) showResult(doc)

It’s a reasonable thing to write. You run some queries, watch the scores go by, and pick a number that seems to separate the good hits from the noise. On nomic-embed-text, 0.75 is a defensible choice — 5 of my 8 queries clear it.

Now swap the model. Same code, same threshold, same bookmarks.

0 of 8. Every single correct answer falls below the cutoff and gets filtered out.

The search doesn’t error. The model loads fine. The ranking is still perfect. The results are just gone — the user types a query and sees an empty list, and nothing anywhere says why.

This is the same shape as a problem I hit tuning GPU memory offload: no crash, no warning, just quietly wrong. Those are the expensive ones.

What I do instead

Rank, don’t threshold. Take the top N by similarity and show them. N is a property of your UI — how many rows fit — and it survives a model swap untouched.

Where I do combine scores, I normalise first:

final = 0.65 × semantic + 0.35 × fuzzy     // each min-max normalised

Min-max normalisation rescales whatever the model handed me into 0–1 within the current result set. The absolute scale stops mattering, which is exactly what I want.

The fuzzy half is there for a different reason: pure semantic search is bad at exact words. When I type a word that’s literally in the title, I want that bookmark first, not something the model considers thematically adjacent.

Don’t do these

  • Don’t hardcode a similarity threshold. It’s model-specific, and nothing warns you when it stops applying.
  • Don’t compare similarity scores across models. 0.70 from one is not 0.70 from another. They aren’t the same unit.
  • Don’t assume a model swap is a drop-in. Dimensions change, so your cached vectors are invalid, and the scale changes, so your tuning is invalid.
  • Don’t trust your memory of a bug. I nearly published a confident claim that my own test then contradicted.

Honest limits

  • This is a small test. 24 documents, 8 queries, short titles. A bigger corpus or longer text could behave differently.
  • Two models only. I didn’t test any of the dozens of other embedding models out there.
  • Neither model is bad here. Both got every answer right. The problem is the number attached to the answer, not the answer.
  • I still don’t know what I actually saw months ago. Model versions move, and I didn’t keep the data. That’s on me.

The thing this came out of

All of this lives in MarkJump, a Chrome extension I built for searching bookmarks. Fuzzy search is the default and works offline with no setup. The semantic search above is optional and off unless you turn it on, and when you do, it only talks to a local Ollama endpoint — localhost, enforced by the host permission itself. No accounts, no telemetry, no cloud.

It’s free, and I’d be glad if you gave it a try: MarkJump on the Chrome Web Store.

Closing

I went looking for a model that was broken, and what I found instead was a number that meant nothing on its own. Both models answered every question correctly. A single hardcoded constant was ready to throw all of it away.

Similar Posts