← all articles

The best vector databases for RAG in 2026

This list is for people who are building retrieval augmented generation into a real product and need to pick where the vectors live. i run small operator-style projects out of Singapore, so my bias is toward things one person can deploy, monitor and pay for without a platform team. if you are at a company with a dedicated data infrastructure group, some of my advice will feel light, but the trade-offs are the same.

Most RAG failures i have seen are not the database’s fault. they come from bad chunking, weak embeddings or no evaluation. i wrote about that in why your RAG answers are wrong, and it is worth reading before you spend a week migrating between vector stores. that said, the database does shape your cost, your latency and how painful filtering and updates are, so the choice is not nothing.

I looked at seven options: two managed-first services, three open-source engines with hosted versions, one Postgres extension and one search engine you may already run. prices change often, so treat the figures below as a snapshot from late September 2026 and check each vendor’s pricing page before you commit.

how I picked

  • retrieval quality features: hybrid search (dense plus keyword), reranking hooks and metadata filtering that does not tank recall. RAG over real documents needs all three.
  • operational burden: how long from zero to a working index, and how much babysitting it needs afterwards. i weighted this heavily because i am the on-call person.
  • cost at small and medium scale: what a few million vectors costs per month, and whether there is a free tier that is usable for a prototype rather than a demo.
  • ecosystem fit: client libraries, integrations with common orchestration code, and how easy it is to leave. see my thinking on lock-in in comparing agent frameworks by what they lock you into.
  • deployment options: self-host, managed cloud, or both, plus where the data can sit. this matters if you have customers who care about residency. The Privacy Wire covers that side better than i can.
  • maturity: how long the project has been around and how clear the docs are about limits.

the picks

Qdrant

Qdrant is the one i reach for first when i want a dedicated vector database. it is written in Rust, ships as a single binary or Docker image, and the payload filtering is genuinely good. filters are applied during the graph search rather than after it, which keeps recall steady when you filter on tenant, date or document type. that pattern is common in RAG, where every query is scoped to a user.

It also supports sparse vectors and hybrid queries in one request, so you can combine dense embeddings with keyword-style sparse scoring without stitching two systems together. the Qdrant documentation is clear and the quickstart works as written, which is rarer than it should be. Qdrant Cloud has a small free cluster for prototypes.

  • pros: strong filtered search, hybrid queries in a single call, easy self-hosting
  • pros: quantization options that cut memory use a lot on larger collections
  • pros: free cloud tier that is enough to build a real proof of concept
  • cons: smaller ecosystem and fewer tutorials than Pinecone or Postgres
  • cons: you own the tuning of quantization and HNSW settings once you grow

Pricing: open source and free to self-host, free small cloud cluster, paid cloud is usage based (check qdrant.tech/pricing for current rates).

pgvector (Postgres)

If your app already has Postgres, start here and only leave when you have a measured reason. pgvector adds vector columns and HNSW and IVFFlat indexes to Postgres. your embeddings sit next to your users, documents and permissions, so joins and transactions just work. deleting a document and its chunks in one transaction is a small thing that saves a lot of bugs.

The catch is scale and tuning. index builds are memory hungry, recall depends on settings like ef_search, and filtered queries can return fewer results than you asked for if the filter is selective. it is fine for hundreds of thousands to a few million vectors on a sensible instance. beyond that i would test carefully. most managed Postgres providers support the extension, so you rarely need to install anything yourself.

  • pros: one database for everything, transactional consistency with your app data
  • pros: no new infrastructure or vendor if you already run Postgres
  • pros: mature tooling for backups, replicas and access control
  • cons: filtered ANN queries need care and can hurt recall
  • cons: heavy index builds and memory use compete with your normal workload

Pricing: the extension is free and open source. you pay for the Postgres instance you already run.

Pinecone

Pinecone is the managed-only option and the easiest to get running when you do not want to think about infrastructure. serverless indexes mean you do not size pods, and the API is small. for a team that wants retrieval to be someone else’s problem, that is a real feature. the Pinecone docs cover namespaces, metadata filters and integrated embedding and reranking options.

The trade-off is that it is a closed hosted service. you cannot self-host it, so the exit path is re-indexing elsewhere, which is fine if you keep your source documents and embedding pipeline separate. costs are usage based, and there is a monthly minimum on paid plans, so it is a poor fit for tiny hobby projects that outgrow the free tier. for anything with real traffic i find the cost predictable enough.

  • pros: no infrastructure to run, fast to first working query
  • pros: namespaces make multi-tenant RAG straightforward
  • pros: good docs and broad integration support
  • cons: managed only, no self-host option
  • cons: paid plans have a minimum spend that is awkward for small side projects

Pricing: free starter tier, paid plans usage based with a monthly minimum (see pinecone.io/pricing for current numbers).

Weaviate

Weaviate sits between a database and a retrieval framework. it stores objects with their vectors, has built-in hybrid search that fuses BM25 and vector scores, and supports modules that can call embedding models for you. if you like the idea of the database handling vectorization at ingest time, that is convenient. the Weaviate developer docs are thorough.

It has more concepts to learn than Qdrant or pgvector: schemas, collections, multi-tenancy settings, vectorizer modules. for a simple RAG app that is more surface than you need. for a larger one with several object types and cross references, that structure earns its keep. it runs self-hosted or on Weaviate Cloud.

  • pros: hybrid search is a first-class feature, not an add-on
  • pros: multi-tenancy support built for SaaS-style RAG
  • pros: self-host or managed, so you can start local and move later
  • cons: more concepts and configuration than the simpler options
  • cons: self-hosted clusters take real memory and some operational care

Pricing: open source to self-host, Weaviate Cloud is usage based with a trial sandbox (check weaviate.io/pricing).

Milvus and Zilliz Cloud

Milvus is the option built for very large collections. it separates storage, query and index nodes, supports several index types including disk-based ones, and is designed to scale to billions of vectors. if you are doing RAG over a huge corpus, that architecture matters. Zilliz is the company behind it and sells the managed version as Zilliz Cloud. the Milvus docs explain the deployment modes.

For a small project it is overkill. the full distributed deployment has a lot of moving parts. Milvus Lite and the standalone mode are easier, and Zilliz Cloud removes the operational side entirely, so i would only self-host the cluster if you have a reason and someone to run it.

  • pros: proven at very large scale, many index types
  • pros: managed option through Zilliz Cloud avoids cluster operations
  • pros: good hybrid and multi-vector support
  • cons: distributed self-hosting is heavy for small teams
  • cons: more to learn than a single-binary database

Pricing: open source and free to self-host, Zilliz Cloud has a free tier and paid usage-based plans (see zilliz.com/pricing).

Chroma

Chroma is the friendly one. pip install chromadb and you have a local vector store in a few lines, which is why it shows up in so many tutorials. for prototyping a RAG pipeline on a laptop it is hard to beat, and it lets you test retrieval ideas before you decide on infrastructure. the Chroma docs are short and readable.

I would not pick the embedded version for production traffic with many concurrent writers. there is a client-server mode and a hosted Chroma Cloud, and both are improving, but it has a shorter track record than the others here. treat it as the prototype tool that can grow into a small production one, and test your own load before you trust it.

  • pros: fastest setup of anything on this list
  • pros: simple API that is easy to teach to a new teammate
  • pros: works fully local with no account
  • cons: shorter production track record than older engines
  • cons: fewer options for tuning and scaling at large sizes

Pricing: open source and free locally, Chroma Cloud is usage based (check trychroma.com for current pricing).

OpenSearch

OpenSearch is on the list for a specific reader: the one who already runs OpenSearch or Elasticsearch for site search or logs. both support dense vector search alongside their strong keyword scoring, so hybrid retrieval is natural. the OpenSearch k-NN documentation covers engines and settings. one system for full text and vectors means one thing to monitor.

The downside is that these clusters are not light. JVM tuning, shard planning and index management are a job. if you have never run a search cluster, do not start here for RAG alone. if you already have one and your documents are already indexed, adding a vector field is often the cheapest path.

  • pros: best-in-class keyword search combined with vector search
  • pros: reuses infrastructure and skills you may already have
  • pros: mature security, access control and observability
  • cons: heavy to operate if you are starting from nothing
  • cons: vector search tuning is less friendly than in dedicated engines

Pricing: OpenSearch is open source and free to self-host. managed versions such as Amazon OpenSearch Service bill by instance and storage.

comparison table

pick price primary strength primary weakness
Qdrant free self-host, small free cloud, usage-based paid filtered and hybrid search smaller ecosystem
pgvector free extension, pay for Postgres one database for everything tuning and filtered recall
Pinecone free starter, paid with monthly minimum zero infrastructure managed only
Weaviate free self-host, usage-based cloud built-in hybrid and multi-tenancy more concepts to learn
Milvus / Zilliz free self-host, free and paid cloud very large scale heavy to self-host
Chroma free local, usage-based cloud fastest to prototype shorter production record
OpenSearch free self-host, instance-based managed keyword plus vector in one system heavy to operate

how to choose

Start with what you already run. if you have Postgres and fewer than a few million chunks, use pgvector and spend your time on evaluation instead. it is the choice with the fewest new failure modes, and you can always move later because your embeddings are reproducible from your documents. keep the raw text and the embedding model name stored so a migration is a re-index job and not a rebuild.

If you need a dedicated store and you are a small team, choose between Qdrant and Pinecone based on one question: do you want to run it? Qdrant gives you control and a cheap self-hosted path. Pinecone gives you a service and a bill. neither is wrong. i lean Qdrant because i like being able to reproduce production on my laptop, but i understand why teams pay to avoid the pager.

Match the database to your retrieval shape. multi-tenant apps where every query filters on a customer should test filtered recall specifically, because that is where systems differ most. if your corpus is mostly exact terms such as part numbers, legal clauses or error codes, hybrid search is not optional, so weight Weaviate, Qdrant and OpenSearch higher. and remember that the embedding model matters as much as the store, so look at the best open-source embedding models in 2026 before you fix your vector dimensions.

Finally, do not skip testing. build a small set of real questions with known good source passages and measure recall at k on each candidate before deciding. i cover a workable approach in testing an AI feature before you ship it. also be honest about what retrieval can and cannot fix. a bigger window does not replace good retrieval, as i explain in what a context window does not solve. browse the rest of the blog index for more on the surrounding stack.

verdict / top pick

My top pick for most people is pgvector if you already run Postgres, and Qdrant if you do not. that is a boring answer, but it matches what has actually held up for me. pgvector removes an entire system from your architecture, and Qdrant is the dedicated engine i trust most for filtered, hybrid retrieval on a small budget.

Pinecone is the pick if you want to pay to not think about it. Weaviate suits apps that lean on hybrid search and multi-tenancy. Milvus with Zilliz Cloud is for very large corpora. Chroma is for prototypes. OpenSearch is for teams already living in search. whichever you choose, keep your source documents and embedding pipeline portable, measure recall on your own questions, and revisit the decision only when the numbers tell you to. this is not financial or purchasing advice, and vendor pricing can change, so verify before you sign up.

Written by Xavier Fok

disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-29.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →