Weaviate vs Qdrant for RAG: Which Vector Database to Pick
A million document chunks embedded with OpenAI’s text-embedding-3-small is a million lists of 1,536 floats. At 4 bytes a float that is about 6.1 GB of raw vectors, before any index overhead. At 10 million chunks it is 61 GB. That arithmetic decides most of the Weaviate versus Qdrant question, because it tells you whether you need a big-memory machine, a quantization plan, or a managed plan billed by size. Which one “searches better” comes a distant second.
Both are open-source vector databases you can run yourself or rent as a cloud service. Weaviate is written in Go and does a lot inside the database: BM25 keyword search, hybrid ranking, embedding modules, and a generative search call that retrieves and prompts an LLM in one request. Qdrant is written in Rust, has a narrower scope, and is very good at filtering vectors by metadata while it searches. My verdicts in short: Qdrant for most single-team RAG projects that start small and self-host, Weaviate for multi-tenant SaaS and for teams that want the database to carry more of the pipeline. If you have under a million chunks and Postgres already, pick neither yet and read pinecone vs pgvector for RAG.
Two caveats. Some axes below, like context window and price per million tokens, belong to language models rather than databases, so I answer what they mean here and say “not applicable” where that is the honest answer. And I have not run both on identical hardware against my own corpus. Speed numbers are vendor-published and labelled that way. I’m a Singapore-based operator running small projects, not a platform team, so this is a view from the cost-and-upkeep side. It is one of several infra comparisons on the blog.
TL;DR comparison table
| axis | Weaviate | Qdrant |
|---|---|---|
| License and language | BSD-3-Clause, Go | Apache 2.0, Rust |
| Cloud pricing model | tiered plans, usage tied to stored vector data | cluster resources (vCPU, RAM, disk), plus a free 1 GB cluster |
| Hybrid search | built-in BM25F plus vector in one query | sparse plus dense vectors fused in the Query API |
| Multi-tenancy | native, one shard per tenant | payload-based tenant partitioning, or custom shard keys |
| Embeddings in the database | vectorizer modules, Weaviate Embeddings | FastEmbed client library, cloud inference |
| Retrieve and generate in one call | yes, generative search | no, you call the LLM yourself |
| Self-host | Docker, Kubernetes, bring-your-own-cloud | single Docker binary, Kubernetes, Hybrid Cloud |
| Support | community forum and Slack, paid support on cloud plans | community Discord, paid support on cloud plans |
| Target user | SaaS builders, teams wanting search and generation in one system | small teams, filter-heavy search, cost-sensitive self-hosters |
Weaviate at a glance
Weaviate grew out of an Amsterdam company started by Bob van Luijt. It is open source under BSD-3-Clause and written in Go. You define a collection with properties, a vector index (HNSW by default, with flat and dynamic options) and, if you want, a vectorizer module that calls an embedding provider for you.
Hybrid search is a first-class query. BM25F keyword scoring and vector scoring are fused in a single call, and the hybrid search docs show how to pick the fusion method and set an alpha weight between the two. Multi-tenancy is part of the data model, and Weaviate Cloud will run all of it for you.
The trade is scope. You get more built-in behaviour and more configuration to understand before your first query.
Qdrant at a glance
Qdrant is from Berlin, open source under Apache 2.0, and written in Rust. The model is plain: a collection holds points, each point has one or more vectors plus a JSON payload, and a search can filter on payload fields during the HNSW traversal instead of after it. That is why people reach for it when chunks carry permissions, dates or source types.
Sparse vectors, multivectors for late-interaction models like ColBERT, and scalar, product and binary quantization are all in the engine. The hybrid queries docs cover running dense and sparse searches side by side and fusing the results. Qdrant Cloud lists a free 1 GB cluster on its pricing page, which is how most people first try it.
Head-to-head
Model quality and benchmarks
Neither database is a model, so there is no quality score. What you can measure is retrieval quality: how often the right chunk lands in your top 5 or top 10. That depends mostly on your embedding model and your chunking. Both use HNSW as the main index, so at matched settings recall lands in roughly the same range. The differences come from quantization defaults, filtering behaviour and how hybrid results are fused.
Qdrant publishes benchmarks with an open-source harness, and Weaviate publishes its own. Each set tends to favour its author, and I haven’t reproduced either. Read them for method, not for a winner.
I’d follow the habit in picking a model by its failure mode, not its benchmark score. Decide which retrieval miss you can live with, exact terms like SKUs and error codes or fuzzy meaning, then test for that. Exact-term misses are where Weaviate’s built-in BM25 helps out of the box. Qdrant gets there with sparse vectors, which is more setup.
Context window
Not applicable. The context window belongs to the LLM you send the chunks to. The database only affects how well you fill it: how many chunks you retrieve, whether you can group hits by document so five chunks from one page don’t crowd out everything else (both have grouping queries), and whether payloads come back with each hit so you skip a second lookup.
Chunk size times top-k has to fit under the model’s limit with room left for the answer. If you’re hitting that wall, what a token limit error actually tells you is a quick read. Weaviate’s generative search builds the prompt for you, so check what it actually sends before you trust your own token maths.
Latency and throughput
Vector search is usually a small slice of a RAG request. As far as I can read the vendors’ published runs, both land at single-digit to tens of milliseconds at moderate scale, while the LLM call takes seconds. Why your LLM app feels slow walks through where the time goes, and the database is rarely the first suspect.
It matters more under heavy filters and at scale. Qdrant’s benchmark page claims the lead on requests per second and latency in most of its scenarios. It is Qdrant’s own test, so I hold it loosely. The Rust versus Go argument comes up a lot here. I wouldn’t decide on it, because memory layout and quantization settings move latency more than the language does.
One physical fact for anyone in Singapore: a cluster in a US region adds well over 150 ms of round trip before any search work happens. Check which regions each vendor offers near you before comparing anything else.
Pricing per million tokens
Neither charges per million tokens for the database itself. Cost arrives in other ways.
Cloud plans count different things. Qdrant Cloud bills for cluster resources (vCPU, RAM and disk), so you can estimate a bill from the arithmetic at the top of this post. Weaviate Cloud sells tiered plans with usage tied to how much vector data you store. I haven’t re-checked either rate card this week and both change, so use the vendors’ pricing pages for current figures.
Compression is the lever. That 6.1 GB of float32 vectors drops to about 1.5 GB with int8 scalar quantization and about 190 MB with binary quantization, before graph overhead. Both engines support both. You pay for it in recall, and you claw most of it back by rescoring the top hits against the original vectors.
Token meters only appear if you use hosted embeddings. Weaviate Embeddings and Qdrant Cloud’s inference option both embed for you, and that’s where a per-token line can show up, so compare it with calling OpenAI or Cohere directly. The bigger line in most RAG apps is generation, not storage. What an AI agent actually costs to run breaks that down.
API ergonomics and SDK quality
Weaviate’s v4 Python client is typed, works with collection objects and talks gRPC underneath. It’s pleasant once the collection config makes sense, but that config surface is large: vectorizers, generative modules, named vectors, index types, tenants. The one-call generative search is the best prototype convenience in either product. It also puts your LLM prompt inside a database request. I would not ship that. If you want control over retries, prompt structure or splitting the work into steps, when to split one prompt into two calls is the reason to keep generation out of the database.
Qdrant’s surface is smaller: create a collection, upsert points, query. The Python client has a local mode that runs in memory or on disk with no server, so tests skip Docker, and the server ships a web dashboard on port 6333. Clients exist for Python, JavaScript and TypeScript, Rust, Go, .NET and Java. Weaviate covers Python, TypeScript, Go and Java. Qdrant gets me to a working query faster. Weaviate goes further before I write glue code.
Self-host vs managed
Qdrant is one Docker container running one binary, and a single node goes a long way. It can keep full vectors on disk through memory-mapped files and hold only the quantized copies in RAM. I’d start there for anything under tens of millions of chunks. Weaviate also runs from Docker in development and Helm on Kubernetes in production. In my reading of the docs it wants more memory for the same vector count, because the HNSW index lives in RAM unless you compress it. Both support sharding and replication once you outgrow one node.
Managed, Qdrant Cloud runs on AWS, GCP and Azure, and Hybrid Cloud puts the data plane in your own Kubernetes cluster with Qdrant managing it. Weaviate Cloud has serverless and dedicated options plus bring-your-own-cloud. The free 1 GB Qdrant cluster is the easiest way to test a managed setup. As far as I know Weaviate has offered time-limited sandbox clusters rather than a permanent free one, so check that before planning around it.
Self-hosting the vector store is one decision. Self-hosting the embedding and generation side is another, and vLLM vs TGI for self-hosted inference covers it.
Data retention and training policy
A vector database is not a model provider, and I’m not aware of either vendor training models on customer vectors. I can’t verify the terms for you and they change, so read the terms and data processing addendum for whichever cloud plan you buy. Self-hosting removes the question, since the data stays on your disk.
Three practical points:
- deletes: both mark records as deleted first and reclaim space later (tombstones and cleanup in Weaviate, segment optimizers in Qdrant), so “deleted” is not instantly “gone from disk”. Ask about backups and snapshots too.
- region: where the cluster sits decides which rules apply to you. This is not legal advice. If you store personal data from Singapore or the EU, check the PDPA and GDPR with someone qualified.
- embedding provider: the chunk text goes to whoever embeds it, whether that is OpenAI, Cohere or a local model. That path is often the bigger privacy decision.
Anything users can write into the index also ends up inside a prompt later, which is why prompt injection in practice applies to a RAG index more than to most apps.
Ecosystem and integrations
Both are first-class in LangChain, LlamaIndex and Haystack, so the framework won’t decide it. Weaviate’s own strength is its module system: vectorizers for OpenAI, Cohere and Hugging Face, rerankers, and generative modules for the one-call flow. Qdrant’s contribution is FastEmbed, a Python library that runs ONNX embedding models on CPU and also produces sparse and late-interaction embeddings. There is also an official MCP server for wiring Qdrant into agent tools, as of my last check.
Use-case verdicts
Metadata-heavy RAG on one machine: Qdrant
Your chunks are tagged with customer, permission group, date and document type, and every query filters on two or three of them. Qdrant filters during graph traversal, a single node with quantized vectors in RAM handles a lot, and it starts from one docker run command. Weaviate filters well too. The difference is how much setup you carry to get there.
Multi-tenant SaaS: Weaviate
Each customer’s data must stay separate, and you may have thousands of customers, most of them idle. Weaviate gives each tenant its own shard and can deactivate or offload the quiet ones. Qdrant’s usual pattern is one collection with an indexed tenant field in the payload, and custom shard keys narrow the gap. That works and it is efficient, but isolation rests on the filter being in every query. A missed filter is a data leak, and I sleep better with isolation in the data model. Narrow win.
Support bot over docs full of exact terms: Weaviate, narrowly
Help centre articles packed with error codes, plan names and version numbers punish pure vector search. Weaviate’s BM25 plus vector query works after one config step. Qdrant reaches the same quality with sparse vectors and fusion, and its ceiling is as high. It just takes more setup. If the help centre is public, those pages are also what search engines and AI crawlers read, and the SEO side of that lives on my sister site, The SEO Desk.
Side project or cheap first production: Qdrant
A free cluster, a local mode for tests, one binary for a box you already pay for. Do the arithmetic first though. At 384 dimensions the free 1 GB holds roughly a couple hundred thousand float32 vectors. At 1,536 dimensions you’ll want quantization.
Who should pick Weaviate
- you’re building multi-tenant SaaS and want tenant isolation in the data model
- you want BM25 and vector search in one query with little wiring
- you like the database calling the embedding provider for you
- you have someone who will read the config docs properly, and a memory budget to match
Who should pick Qdrant
- one small team, one box, and you want a working index this afternoon
- your queries filter on metadata heavily
- cost is tight and a permanent free cluster or a cheap VPS matters
- you’d rather keep the LLM call in your own code
Be aware that you assemble hybrid search yourself with sparse vectors. I haven’t had to run that at scale, so I can’t tell you where it hurts.
Verdict overall
The frontmatter says “it depends” and I mean it. If someone made me pick one for a first RAG project, I’d pick Qdrant: less to configure, a lower cost of entry, and a permanent free cluster to test on. If the project is multi-tenant SaaS, I’d pick Weaviate.
Before believing any benchmark page, including the ones I linked, take 500 of your own chunks and 20 real questions and run both. Evaluating an AI tool in an afternoon is how I’d set that up. And if pgvector turns out to do the job, stop there.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-09-19.