← all articles

Paid embeddings vs open source embeddings

I priced out two ways to add semantic search to a small tool last week: call OpenAI’s embedding API and pay per token, or run an open model on a spare box I already had spinning in the rack. Both work. Both return the same kind of thing, a list of numbers that represents what a piece of text means. The difference is who runs the model and who eats the cost, and that decision matters more than most tutorials let on.

If you’ve read anything about retrieval-augmented generation, semantic search, or recommendation systems, you’ve run into embeddings without necessarily being told what they are. This is the plain version.

what it is

An embedding is a fixed-length list of numbers, usually somewhere between 256 and 3072 of them, that a model produces from a piece of text (or an image, or audio, but text is the common case here). Two pieces of text that mean similar things end up with vectors that sit close together in that numeric space. Two pieces that mean different things end up far apart. That’s the whole trick. Search, clustering, deduplication and RAG all lean on that one property.

Paid embeddings means you send your text to someone else’s API over HTTPS and they hand you back the vector, billed per token. OpenAI’s text-embedding-3-small and text-embedding-3-large are the two most people reach for first, and Cohere and Google both sell equivalents. Open source embeddings means you download the model weights yourself, usually from Hugging Face, and run the inference on your own hardware. No per-call fee, but you’re now responsible for the box it runs on.

how it works

Under the hood almost all of these models are transformer encoders, the same family of architecture behind the LLMs everyone talks about, just trained with a different objective. Instead of predicting the next word, an embedding model is trained, often with contrastive learning, where it sees pairs of similar and dissimilar text and learns to pull the similar pairs together in vector space, to produce a representation that captures meaning rather than exact wording. That’s why “cheap flights to Bali” and “affordable Bali airfare” end up close together even though they don’t share a single distinctive word.

Once you have vectors, you need somewhere to put them. That’s a vector database, and I’ve written separately about how the main options compare if you want the fuller breakdown. At query time you embed the search text with the exact same model that embedded your documents, then find the nearest neighbors by cosine similarity or dot product. But that “exact same model” part is the detail that trips people up, more on that below.

One thing OpenAI did with the text-embedding-3 family is worth knowing: the vectors support truncation. You can cut a 3072-dimension embedding down to a shorter length and lose only a little retrieval accuracy, a trick that comes from Matryoshka representation learning, where the model is trained so the important information sits in the first dimensions.

Handy if you’re paying for vector storage by the byte.

why it matters

A few reasons this choice actually changes what you build, not just what you pay:

  • cost crosses over at volume. OpenAI’s text-embedding-3-small runs at $0.02 per million tokens according to OpenAI’s own pricing docs, which is close to nothing for a prototype. Embed a few hundred million documents a month and it adds up fast enough that a GPU you already own starts looking free by comparison.
  • data residency. if you’re embedding customer support tickets, medical notes, or anything with a privacy obligation attached, sending that text to a third party API is a decision your legal or compliance person should sign off on, not something you default into because the docs made it a one-liner. self-hosting keeps the text on hardware you control. The Privacy Wire has more on data residency tradeoffs generally if that’s the part you’re weighing.
  • vendor drift. providers update embedding models. if a provider ships a new default version and you re-embed only new documents against it, your index now has vectors from two incompatible models sitting side by side, quietly returning worse search results with no error message anywhere.
  • offline and edge cases. a self-hosted model doesn’t need a network round trip, which matters if you’re building something that has to work on a flaky connection, or you just don’t want a hard dependency on an API staying up.

common misconceptions

open source means worse quality. not true, at least not automatically. several open models sit at or near the top of the MTEB leaderboard on Hugging Face, including checkpoints from BAAI and Alibaba’s GTE family, ahead of some paid options on the same retrieval benchmarks. I haven’t run every model on that board against my own data, so treat it as a starting point for testing, not a final verdict.

bigger dimension count means better results. it means more storage and more compute in your vector index, not automatically better retrieval for your specific task. a well-trained 768-dimension model can beat a poorly fitted 3072-dimension one on your actual documents.

you can mix and match embedding models in one index. you can’t, not without re-embedding everything. vectors from different models, or even different versions of the same model, live in different coordinate spaces, and comparing them returns garbage that still looks like a valid similarity score. if you think you might switch providers later, it’s worth reading how to structure the provider-switching layer before you’ve got ten million vectors to migrate.

free means zero cost. self-hosting trades the per-token bill for GPU time, a serving stack, and your own hours when something breaks at 2am. if you’re going to run models locally day to day, Ollama and LM Studio are the two tools worth knowing first, and they’ll teach you fast whether self-hosting is actually less hassle for your setup or just cheaper on paper.

where to go from here

A few places to go next if this is the start of a bigger decision, not the whole of it:

My honest take, for what it’s worth: unless you’re already running your own GPUs for something else, or you have a real reason to keep the text off someone else’s servers, start with a paid API. the few dollars a month you’d save self-hosting a small model rarely covers the time you’ll spend keeping it fed and updated. scale changes that math, but most projects never get there.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-14.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →