← all articles

The best open source embedding models in 2026

A 4096-dimension float32 vector takes 16,384 bytes. Multiply that by 10 million chunks and you are holding about 164GB of vectors before the index adds anything on top. The same corpus at 384 dimensions is about 15GB. That gap often matters more than a two-point difference on a leaderboard.

This list is for people who run their own retrieval. Search over docs, the RAG layer behind a support bot, duplicate detection, clustering, semantic matching inside a product. You want weights you can download, a license you can read in one sitting, and the option to move onto your own hardware if an API bill or a data policy forces the issue. I write from Singapore, where English mixed with Mandarin, Malay and Tamil is ordinary, so I weigh multilingual behavior more heavily than a US-only writer would.

I picked eight models. One carries a license caveat, and I flag it where it matters. I have not re-run all eight on one shared corpus, so the few scores I quote are vendor-reported and dated. Use them to build a shortlist, then test on 200 of your own real queries. That beats any leaderboard, including MTEB. New models land every month, so anything released recently may be missing here.

How I picked

  • license: Apache 2.0 and MIT count as open source here. Gemma’s terms are open weights with a use policy attached, and I flag that where it comes up. I left out jina-embeddings-v3 because its weights are CC BY-NC 4.0, which rules out most commercial use.
  • quality: retrieval scores on MTEB and its multilingual variants, as reported on the model card or in the vendor’s release post.
  • context length: 512 tokens forces small chunks, 8192 lets you embed a whole page, and that changes how you build the pipeline.
  • footprint: parameter count, output dimensions, and whether Matryoshka truncation lets you shrink vectors later without retraining.
  • languages: a model that is great in English and shaky in Malay is a different product for a Singapore shop than for a US one.
  • tooling: how much friction it takes to load in sentence-transformers or run through Ollama or llama.cpp.

The picks

Alibaba Qwen (Qwen3-Embedding)

Qwen3-Embedding arrived in June 2025 in three sizes, 0.6B, 4B and 8B parameters, all Apache 2.0. Qwen’s release post reported the 8B model first on the MTEB multilingual leaderboard at that time, with a score of 70.58. That is a vendor claim from mid-2025 and the board has moved since, so check it.

What I like is the ladder. The three sizes output 1024, 2560 and 4096 dimensions, all support Matryoshka-style truncation, and all take up to 32k tokens plus an instruction string that tells the model what the query is for. Prototype on the 0.6B, move up later, and the pipeline keeps its shape. If you hand-roll the transformers code, note that it uses last-token pooling. sentence-transformers handles that for you.

  • pro: top-tier multilingual retrieval for an open model, going by Qwen’s own MTEB report
  • pro: 32k context and instruction-aware queries
  • pro: three sizes behind one interface, so moving up is a config change plus a re-embed
  • con: the 8B needs about 16GB for weights alone in bf16, so a 24GB card leaves little room for batching
  • con: 4096-dim vectors put you straight into the 164GB-per-10M problem unless you truncate

Pricing: free weights under Apache 2.0. Your cost is GPU time.

Link: Qwen3-Embedding-8B on Hugging Face

BAAI (BGE-M3)

BGE-M3 comes from the Beijing Academy of Artificial Intelligence, published in early 2024 under the MIT license. It is 568M parameters, outputs 1024 dimensions, takes 8192 tokens and covers 100+ languages. The M3 in the paper stands for multi-linguality, multi-functionality and multi-granularity.

The multi-functionality part is why it stays on my list. One forward pass gives you a dense vector, sparse lexical weights and ColBERT-style multi-vector output. So hybrid search comes from one model instead of a model plus a separate BM25 pipeline. You do need a vector store that accepts sparse vectors, and Qdrant and Milvus both do.

  • pro: dense, sparse and multi-vector outputs from one model
  • pro: MIT license and 8192-token context
  • pro: loads through FlagEmbedding and sentence-transformers, and Ollama lists it as bge-m3
  • con: it dates from early 2024, and newer models such as Qwen3-Embedding report higher multilingual MTEB scores
  • con: multi-vector mode stores one vector per token, so start with dense plus sparse only

Pricing: free, MIT license.

Link: BAAI/bge-m3 on Hugging Face

Snowflake (Arctic Embed 2.0)

Snowflake released arctic-embed-m-v2.0 (305M parameters) and arctic-embed-l-v2.0 (568M) in December 2024 under Apache 2.0. Both are multilingual, both take 8192 tokens, and both were trained so that Matryoshka truncation costs little retrieval quality. The pitch is retrieval quality per stored byte, which matters once you are embedding tens of millions of chunks.

Run the numbers. At 256 dimensions in float32, 10 million vectors take about 10GB. At 1024 dimensions they take about 41GB. Whether the smaller version holds up on your data is something to test, but this family was built for that trade.

  • pro: Apache 2.0, two sizes, 8192-token context
  • pro: Matryoshka training makes 256-dim vectors a serious option
  • pro: multilingual in both sizes
  • con: queries need a prefix and documents do not, and forgetting it costs recall without raising an error, so copy the exact prompt from the model card
  • con: smaller community than BGE or Qwen, so fewer worked examples when something odd shows up

Pricing: free, Apache 2.0.

Link: Snowflake/snowflake-arctic-embed-l-v2.0 on Hugging Face

Nomic AI (Nomic Embed)

Nomic Embed is the pick when you want to know what the model was trained on. With nomic-embed-text-v1 in February 2024, Nomic released the weights under Apache 2.0 and published the training code and training data alongside them. Very few embedding models have done that.

The v1.5 model is 137M parameters, 768 dimensions, English only, with 8192 tokens and Matryoshka truncation down to 64 dimensions. The v2 mixture-of-experts release in early 2025 went multilingual at 475M total parameters, about 305M active, but its context is 512 tokens. So you pick your compromise. Both use task prefixes, search_query: and search_document:. Check the model cards for current details.

  • pro: open training data and code, so you can audit or retrain
  • pro: v1.5 at 137M parameters is light enough for CPU serving at modest volume
  • pro: Matryoshka on v1.5, 768 dimensions down to 64
  • con: v1.5 is English only, and v2 caps context at 512 tokens
  • con: the card asks for trust_remote_code=True, so read that code before production

Pricing: free, Apache 2.0. Nomic also sells a hosted API, and I have not compared its rates.

Link: nomic-ai/nomic-embed-text-v1.5 on Hugging Face

Google (EmbeddingGemma)

EmbeddingGemma launched in September 2025 at 308M parameters, 768 dimensions and a 2048-token context, trained on 100+ languages. Google’s announcement says it runs in under 200MB of RAM when quantized and calls it the highest-ranking open multilingual embedding model under 500M parameters on MTEB at release. Those are Google’s claims, dated September 2025. Matryoshka lets you cut 768 down to 512, 256 or 128.

The license needs its own paragraph. Gemma ships under Google’s Gemma Terms of Use with a prohibited-use policy. That is not Apache 2.0 or MIT, and it is not what the Open Source Initiative would call open source. The terms permit commercial use with conditions, so read them. I am relaxed about it for an internal tool and would want a lawyer’s eyes on it before shipping inside something I resell. This is not legal advice.

  • pro: small enough for phones and laptops, under 200MB of RAM quantized per Google
  • pro: 100+ languages in 308M parameters
  • pro: Matryoshka down to 128 dimensions
  • con: Gemma terms rather than an OSI license, and the Hugging Face download is gated behind accepting them
  • con: 2048-token context means more chunking than the 8192-token models

Pricing: free weights under the Gemma terms.

Link: google/embeddinggemma-300m on Hugging Face

IBM (Granite Embedding)

IBM’s Granite Embedding models are Apache 2.0 and pitched at enterprises. The December 2024 release had English models at 30M and 125M parameters and multilingual ones at 107M and 278M. The R2 English models in 2025 moved to a ModernBERT base and an 8192-token context, at 149M and 47M parameters. The sizes are the reason to look. The 30M English model outputs 384 dimensions, which puts its storage cost in MiniLM territory. Check the cards for current sizes.

  • pro: Apache 2.0 across the whole family
  • pro: sizes from 30M to 278M, small enough for CPU serving
  • pro: R2 English models take 8192 tokens
  • con: the 278M multilingual model officially covers 12 languages, including Chinese and Japanese but not Malay or Tamil
  • con: less independent testing to lean on than for BGE or E5

Pricing: free, Apache 2.0.

Link: ibm-granite/granite-embedding-278m-multilingual on Hugging Face

intfloat (multilingual-e5)

multilingual-e5 was written by Microsoft researchers and published under the intfloat account on Hugging Face, MIT licensed. The model I would try is multilingual-e5-large-instruct: 560M parameters, 1024 dimensions, 512-token limit, about 100 languages. It is old by 2026 standards. It stays on the list because its XLM-RoBERTa base covers Malay, Tamil, Indonesian, Thai and Vietnamese, which matters to me more than a few benchmark points. The model card itself warns that low-resource languages may see degraded performance, so test them.

The instruct variant wants a one-line task description on queries, in the form Instruct: {task} then Query: {query}. The plain variants use query: and passage: prefixes. Get either wrong and recall drops quietly.

  • pro: MIT license and about 100 languages, including regional ones
  • pro: mature, with well-documented failure modes
  • pro: 1024 dimensions and wide tooling support
  • con: 512-token limit means small chunks
  • con: older architecture than the newest picks

Pricing: free, MIT license.

Link: intfloat/multilingual-e5-large-instruct on Hugging Face

sentence-transformers (all-MiniLM-L6-v2)

all-MiniLM-L6-v2 is 22.7M parameters, 384 dimensions, about a 90MB download, Apache 2.0, English. Its model card says it was trained on a dataset of 1 billion sentence pairs, and it caps input at 256 word pieces. It is the default in most tutorials, so people assume it is obsolete. On modern retrieval benchmarks it does sit well below the models above. But for FAQ matching, deduplication, tagging or a laptop prototype the gap often does not matter, and I wrote about the general argument in when a small model is the right choice.

  • pro: 384 dimensions, about 15GB per 10 million float32 vectors
  • pro: runs on CPU, no GPU needed for modest volume
  • pro: every vector store and framework has an example for it
  • con: the 256 word-piece limit truncates long chunks without warning
  • con: English only, with older training data

Pricing: free, Apache 2.0.

Link: sentence-transformers/all-MiniLM-L6-v2 on Hugging Face

Comparison table

Pick Price Primary strength Primary weakness
Qwen3-Embedding Free, Apache 2.0, GPU for 4B and 8B Top multilingual quality, 32k context 8B needs about 16GB for weights
BGE-M3 Free, MIT Dense, sparse and multi-vector in one model Older, 568M parameters
Arctic Embed 2.0 Free, Apache 2.0 Vectors that survive truncation Query prefix, smaller community
Nomic Embed Free, Apache 2.0 Open data and code English only or 512 tokens
EmbeddingGemma Free, Gemma terms Tiny footprint, on-device Not an OSI license
Granite Embedding Free, Apache 2.0 Small sizes, 8192 tokens on R2 12 official languages
multilingual-e5 Free, MIT Wide language coverage 512-token limit
all-MiniLM-L6-v2 Free, Apache 2.0 Tiny and CPU friendly 256 word pieces, English only

How to choose

Start with your language mix and your chunk length, in that order. English only with short chunks: MiniLM, Nomic v1.5 or a Granite English model will do, and you save real money on storage. Mixed languages: Qwen3-Embedding, BGE-M3 or multilingual-e5, and test your actual languages. “Supports 100 languages” only tells you the base model saw them. Retrieval quality in each one is a separate question.

Then decide whether to self-host at all. OpenAI priced text-embedding-3-small at $0.02 per million tokens when it launched in January 2024, and its pricing page has the current figure. At that launch price, embedding 10 million chunks of 500 tokens each is 5 billion tokens, about $100 once. If a one-off index build is your whole workload, a rented GPU will struggle to beat that. Self-hosting wins when you re-embed often, query volume is high, or the text cannot leave your network.

Vectors from different models cannot be compared. Changing models therefore means re-embedding the whole corpus, and that is the lock-in people notice last. Matryoshka truncation shrinks what you store. It does nothing for the re-embed. I wrote about the wider cost of that kind of move in what changes when you move from one provider to another. And a bigger context window on your generation model does not remove the need to retrieve well, which I covered in what a context window does not solve.

Test before you commit. Take 200 real queries, label the chunks that should come back, and measure recall@10 for two or three candidates. That is an afternoon. Here is the opinion you can argue with: skip the 8B for most projects. Start at the 0.6B or BGE-M3 and move up only if your own recall numbers justify the storage bill. Once something is live, log retrieval quality, and the best LLM observability tools in 2026 covers tooling for that. If you use embeddings for SEO work such as clustering keywords or finding near-duplicate pages, the seo desk blog is our sister site for that, and the rest of what I have published is in the blog index.

Verdict / top pick

Top pick: Qwen3-Embedding, starting at the 0.6B size. It gives the best mix of multilingual quality, 32k context and Apache 2.0 terms in this group, going by the vendor reports, and the same interface scales to the 4B and 8B if your own tests show a gap worth paying for.

BGE-M3 is the runner-up and the better choice if hybrid search is the goal, because one model gives you dense and sparse vectors. all-MiniLM-L6-v2 is the pick for a CPU-only box or a quick English prototype. EmbeddingGemma is the one for anything that has to run on a device, once you are comfortable with the Gemma terms.

My ranking rests on model cards and vendor reports, because I have not tested all eight on the same data. Run your own 200 queries before you trust it.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-27.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →