← all articles

How to choose an embedding model

embeddings rag retrieval model-selection

Ninety days notice. That was the email, and the model it was retiring held every vector in an index I had been adding to for the better part of a year.

Replacing it meant embedding the whole corpus again from scratch, because a vector from the new model and a vector from the old one cannot sit in the same index and be compared to each other. Nobody puts that on a model card. It is the single fact that should decide which model you pick.

Picking this is closer to picking a column type

An embedding model is the only component in a retrieval stack whose output you keep permanently.

The chunker you can change on a Tuesday and reindex whatever you feel like. The reranker swaps in an afternoon. Even the vector store can be migrated, because the vectors themselves survive the move intact.

Change the embedding model and every number you have stored becomes meaningless at once. All of them, back to the first document you ever ingested.

So what follows is weighted towards the parts you cannot undo, and away from quality scores, which move far less between reasonable models than the marketing implies.

Dimensions are a commitment

Vector length is the first line on every model card and the last thing anyone reads. Common options run 384, 768, 1024, 1536 and 3072. The spread across that range is eight times, and it multiplies four separate things:

  • bytes held on disk and in memory for every vector you keep
  • bytes read per query, which is what your latency actually tracks
  • arithmetic per distance comparison
  • index build time, every time you rebuild

At 20,000 chunks, none of that is worth an argument. It is a rounding error on a laptop. Pick something reasonable and move on.

At 20 million it is the design.

The pattern I keep running into is a team at the first number choosing as though they were at the second, buying the biggest vector on offer for a corpus that would fit in a spreadsheet. The reverse is rarer and worse, because you discover it at the exact point where you can least afford a migration.

So work out what your chunk count will be in a year. That is the real input here, and almost nobody calculates it before installing anything.

One property is worth hunting for specifically. Some models are trained so the vector can be truncated and still work, with the useful signal packed into the leading dimensions. Keep the first 512 numbers of a 1536 vector and you keep most of the retrieval quality.

That is the only reversible knob in the whole decision. I would trade a couple of benchmark points for it without thinking twice.

The limit that truncates without telling you

Every model has a maximum input length measured in tokens. Go past it and nothing raises. The text gets cut and you get a perfectly well formed vector for the first part of your chunk.

Two things make this nastier than it sounds.

The failure is invisible at every layer you might inspect. The chunk on disk is complete. The vector has the right shape and a sensible magnitude. Retrieval returns ranked results with plausible scores. The only symptom is that certain documents never come back for questions they obviously answer, and you will attribute that to something else for weeks.

Then there is the counting. Tokens are not characters. English averages roughly four characters per token, code is denser than that, and text in a script other than Latin can be two or three times worse again. A chunk you measured in characters and declared safe can be well over the line.

The fix is one assertion. Count tokens with the model’s own tokeniser before you send anything and raise on whatever exceeds the limit, rather than letting a client library silently trim it for you. Ten minutes of work against a bug that hides for months.

Symmetric similarity and retrieval are different jobs

This is the property that does the most damage when you get it wrong, and the one least likely to show up in a leaderboard row.

Some embedding models are trained on pairs of texts that mean the same thing, so two paraphrases land close together. That is symmetric similarity. It is the right tool for deduplicating a support archive or clustering tickets by topic.

Others are trained on question and passage pairs. A nine word question has to land near a three hundred word passage that answers it, and those two texts share very little surface vocabulary. Learning to cross that gap is a different skill from spotting paraphrases.

Point a symmetric model at question to passage retrieval and it works, badly, in a way that is genuinely hard to attribute. Short queries pull short documents. Questions retrieve other questions, so your FAQ page outranks the page carrying the real answer. I lost weeks blaming my chunking for a symptom that was entirely this.

A related trap sits right next to it. Several models expect a prefix on the input, one string for a query and a different one for a document, because that is how the model is told which role the text is playing. Leave them out and nothing complains. Quality just settles a few points below where it should be, permanently.

Both of those live in the model’s usage documentation rather than its benchmark entry. So does the distance measure it was trained against, and whether it wants vectors normalised to unit length before you store them. Read that page. It is usually shorter than the leaderboard.

Multilingual means two different things

If your corpus or your users are not entirely English, this cuts the field before anything else gets a vote.

The weaker claim is that the model handles each language competently on its own. The stronger claim is that a text and its translation land near each other, so a question asked in one language can retrieve a passage written in another. The second is rarer and much harder, and it is what people assume they are buying.

Singapore makes this concrete for me. Plenty of the useful text here is mixed, with English technical vocabulary dropped into another language mid sentence, and tokenisers handle that with varying grace. A coverage list tells you nothing about it. Fifty of your own real sentences will.

Deprecation is a migration on somebody else’s calendar

Back to the email. A hosted model can be withdrawn, and the notice period is fine. The problem is that the replacement has a different vector space, so accepting it means re-embedding everything in a window you did not choose.

Running an open weights model yourself changes the shape of the risk instead of removing it. You have a file. Nobody can retire a file. In exchange you own a GPU, or accept slower inference on CPU, plus the work of operating it.

For a large corpus that changes slowly, I take that trade every time. For twenty thousand chunks it is plainly not worth it, and I would use a hosted model and accept that a forced migration is somewhere on the horizon.

Test it on fifty of your own queries

Fifty real queries, pulled from your logs or your support inbox, typos included. Not fifty questions written by someone reading your own documentation, which produces an exam your system is guaranteed to pass.

Label each one by hand with the chunk that actually answers it. That hour is the price of admission and there is no shortcut through it.

Then embed the corpus once per candidate, run the fifty, and count how often the right chunk lands in the top five.

People skip this because embedding a corpus twice sounds expensive. It is billed per million tokens like everything else, and for a few hundred thousand chunks it costs less than the day you would otherwise spend arguing about it in a design document.

The reason to measure on your own material is narrow: a general benchmark is an accurate measurement of a general corpus, and yours has your vocabulary, your chunk sizes and the particular way your users ask for things.

When I ran it, the order I got did not match the published one. Three models came out close enough that I could not honestly call a winner on retrieval quality, and the decision fell to dimensions and context length instead.

The top of the table is usually the wrong pick

The models up there tend to be the largest. Largest means the longest vectors, the slowest inference and the highest price per million tokens. They earned that rank as an average across a wide spread of tasks, and you are running one task, on one corpus, for one kind of user.

Three or four rows down, at half the dimensions and several times the throughput, is right more often than the ranking suggests.

I have never regretted taking the smaller model. I have regretted the larger one on two projects, and in both cases the reason was printed on the model card before I started.

What I have not tested

Fine-tuning an embedding model on my own pairs. Everything I have read says a few thousand labelled examples from your own domain beats any off the shelf choice comfortably, and it sounds right, and I have not done the work, so I am not going to recommend it.

I also cannot hand you a chunk count where the dimension cost starts to hurt. It depends on your store, your hardware and your query pattern. What I can say is that it arrives sooner than people expect, and an hour with your own index and a stopwatch will tell you more than any comparison table. The per model numbers I keep, dimensions and context limits against current pricing, are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →