← all articles

Vector databases compared: what you actually need in 2026

vector-database rag embeddings retrieval

The second time I built retrieval I deleted the vector database and replaced it with a file. Queries got faster. The answers got better too, though that had nothing to do with the file.

Most comparison pieces open by assuming you need one of these. A good share of readers do not, so the table starts a step earlier than usual.

The verdict

Where you actually are What to use What it costs you
Under ~100k chunks, corpus updates weekly FAISS or hnswlib, index written to disk You rebuild the whole index on every change
Already run Postgres, under a few million vectors pgvector Slower at the top end, and nobody is impressed
Heavy metadata filters, writes every few seconds Qdrant or Milvus, self hosted You now run a stateful service and own its backups
No ops appetite, spiky traffic, two engineers A managed store such as Pinecone Query volume bills you, storage does not
Retrieval quality is your actual complaint None of these. Fix chunking, add a reranker An afternoon, plus some ego

The one number that decides it

How many chunks are you searching over. That is the question, and almost nobody asks it before they start installing things.

Under a hundred thousand, an in memory index answers in single digit milliseconds, costs nothing to run, adds no network hop, and rebuilds from scratch while the kettle boils. I ran exactly that for months. It never once became the bottleneck. What did was my chunking and my queries, and no database was going to fix either.

So build the embarrassingly simple version, measure it, and move when something hurts.

Four things that push you to a real database

Memory first. When the index stops fitting comfortably in the RAM of the box you want to run, you need something that pages from disk sensibly. That line arrives earlier than people expect, because the vectors are only part of the footprint and the graph sitting on top of them is not small.

Then write patterns. A static index is lovely when the corpus updates weekly and miserable when it updates every four seconds.

Then filtering, which is the one that catches people out. You almost never want the nearest neighbours overall. You want the nearest neighbours belonging to this tenant, in this language, published after March. Doing that efficiently instead of post filtering a top 1000 is hard engineering, and the main thing a purpose built store does that a bare index does not.

And the boring one: operations. Several services reading the same collection, access control, backups somebody else has already thought about. That is often the real reason, dressed up as scale because scale reads better in a design doc.

If none of those four describe you, you are picking infrastructure for a problem you do not have.

Do the arithmetic before you believe any vendor

It is one multiplication.

A vector is a list of floats. The length comes from whichever embedding model you picked, and the common ones sit between 384 and 1536 dimensions. Each float is 4 bytes unless you change something.

So a 1536 dimension vector is about 6KB. A million of them is roughly 6GB, before the HNSW graph, which can add a serious fraction again.

Reaching for the largest embedding model available, because bigger sounds better, quadruples that against a 384 dimension model that may score within a point of it on your content. Four seconds of thought, a cost you carry for the project’s life.

There is a cheaper lever underneath. Store those floats as int8 rather than float32 and the footprint drops by four. For retrieval the accuracy loss is usually small, because you are ranking rather than computing anything exact, and you can measure the damage in an afternoon. Almost no tutorial mentions it.

The three categories, and the comparison people get wrong

Libraries give you an index and leave the rest to you. No server, lowest cost, most work.

Self hosted databases hand you real filtering and incremental writes, in exchange for running a stateful service and owning its backups.

Managed services put all of that on somebody else’s pager and bill you monthly or per query.

The mistake I see repeatedly: benchmark a library against a managed service on query latency, declare the library the winner, then spend six months rebuilding half a database around it. Compare on total work, not on the p99.

Postgres is the boring right answer more often than it should be

If you already run Postgres, it does vector search now. pgvector gives you similarity search inside the database you already back up, already monitor, already apply row level security to, and already join against your application data.

That last clause is the quiet advantage. The filtered queries that are hard for a dedicated store are trivial when the vectors live in the same table as the metadata. It is a WHERE clause.

It will lose to a tuned dedicated store at very large scale. Below that, it removes an entire system from the architecture, and removing a system is worth more than shaving milliseconds nobody was measuring.

I would rather explain a slow query than explain why a separate stateful service dropped writes during a deploy.

What the benchmark tables leave out

Every vendor publishes numbers. Read almost all of them as marketing, for three specific reasons.

Recall. Approximate search is approximate, and every system has a dial trading accuracy against speed. A latency figure with no recall figure beside it tells you nothing. 4ms at 90 percent recall and 4ms at 99 percent recall are different products sold under one name.

Filters. Benchmarks run on clean public datasets with uniform vectors and nothing filtered. Your data is lumpy, your queries are filtered, and filtered performance is frequently much worse than the headline. I have never seen a vendor publish that number.

Concurrency. Most benchmarks fire one query at a time. You will fire hundreds. Behaviour under concurrent load is where these systems actually separate, and the least reported figure in the category.

So run your own. Ten thousand of your real chunks. A hundred of your real queries, with your filters switched on. Measure recall next to latency. An afternoon of that beats every comparison table on the internet, including the ones I publish.

Where the bill actually comes from

Managed pricing has a storage line and a query line. People budget the storage line correctly, because stored vectors are easy to count. The query line breaks budgets.

Retrieval augmented systems make several searches per user interaction. Add a reranker pulling 30 candidates, add a follow up retrieval, and one question quietly becomes three or four billable searches.

Multiply traffic by searches per interaction before you compare anything. I have watched a team choose the cheaper provider on the storage line and open an invoice dominated entirely by queries.

Self hosting inverts the curve. You pay for the box busy or idle, which is terrible at low volume and excellent at high steady volume.

The part I got wrong

When my retrieval was bad I assumed the store was at fault. It never was, in either project. Swapping the store moved result quality by a margin I could barely see above noise.

What moved it was how I split documents. Respecting section boundaries instead of cutting every 512 characters. Keeping enough surrounding context that a retrieved fragment still makes sense on its own. Picking an embedding model suited to the content instead of whichever one appeared in the tutorial I copied from.

All three are free. The database is plumbing. Plumbing matters when it leaks, and mine was not leaking.

Two changes that beat any migration

Hybrid search first. Pure vector similarity has a well known weakness: it is excellent at meaning and unreliable at exact strings. Product codes, error identifiers, surnames, version numbers. The embedding smears those into a neighbourhood of similar looking tokens, which is precisely wrong when somebody typed an exact identifier because they wanted that exact thing.

Run BM25 alongside the vector search and merge the two result sets. That fixes a category of failure no migration addresses. If you are on Postgres, full text search is already sitting in the same database.

Then reranking. Pull 20 or 30 candidates instead of five, push them through a cross encoder that scores each against the query properly, keep the best handful.

One extra step, and the largest single quality improvement available in most retrieval stacks. The first search has to be fast across everything, so it approximates relevance cheaply. The reranker only looks at 30 documents, so it can afford to be accurate. Cheap and broad, then expensive and precise.

I put off adding one for months because it felt like more complexity to own. It was worth more than every infrastructure decision I agonised over combined.

Keep the store behind fifty lines of your own code

Whatever you pick, put the write path and the query path behind one small interface you control. Mine is about fifty lines and it has paid for itself twice.

The store is the component most likely to change, and it changes for reasons outside your control. Pricing moves. A project gets acquired. A managed service rewrites its terms with 30 days notice. When that happens you want the swap to be an afternoon rather than a rewrite, and the only thing buying you that is having refused to let a vendor’s client library spread through your codebase.

It also makes comparison cheap. Point the interface at a second backend, replay your real queries, measure recall and latency on your own data. Being able to test alternatives is what stops the first choice mattering much.

Current pricing and the full per system tables are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →