Most vector-database decisions I have reviewed were made in the wrong order. The team read a benchmark, picked a store, built the retrieval pipeline on top of it, and then discovered what their latency budget actually was when the first users complained. By that point the store was load-bearing and the conversation had become about tuning rather than choosing.
The order that works is the reverse. Write the latency budget first. Size the index with arithmetic, not hope. Then let those two numbers pick the store, because in most cases they pick it for you, and the benchmark you were about to read becomes irrelevant.
Step one: the latency budget, written down
A retrieval-augmented request has five or six stages, and only one of them is the vector search. For a chat-style product with a 1,200 ms p95 target to first token, a realistic budget looks like this:
| Stage | Typical p50 | Typical p95 | Notes |
|---|---|---|---|
| Query embedding (hosted API) | 40 ms | 120 ms | Network round trip dominates; self-hosted small models get this to 5–15 ms |
| ANN search (top-k, k=20–50) | 5–20 ms | 30–60 ms | Warm index, in memory |
| Metadata filtering | 0–5 ms | 5–200 ms | Not a separate stage — this is what the filter adds to the search. The wide range is the whole article |
| Reranking (cross-encoder, 20–50 candidates) | 60 ms | 180 ms | Often the largest non-LLM slice |
| Prompt assembly | 2 ms | 5 ms | Negligible unless you are doing something odd |
| LLM time-to-first-token | 300 ms | 700 ms | The budget you do not control |
Those figures are a worked example built from public sizing arithmetic, not a trace from a system I run. Copy the shape, not the numbers. And note that the p95 column sums to more than 1,200 ms. That is the correct and uncomfortable result: stage p95s do not co-occur, so the sum is a worst case rather than the end-to-end p95, but it tells you there is no slack for a stage that misbehaves.
Two things fall out of writing this table. First, the ANN search is around 5 percent of the budget on a good day. A store that is twice as fast at pure vector search saves you single-digit milliseconds at p50 and perhaps thirty at p95. That is not a reason to choose anything. Second, the line with the widest variance is metadata filtering, and it is the one most teams do not think about until the p95 goes wrong. Filtering behaviour, not raw search speed, is the real axis of the pgvector-versus-dedicated decision.
Step two: size the index honestly
The arithmetic is simple and almost nobody does it before choosing.
A 1,536-dimension float32 embedding is 6,144 bytes. Ten million of them is 61 GB of raw vectors. The neighbour lists are not where the memory goes, which is the part people get wrong: at a typical m of 16, each node carries a few dozen four-byte links, well under a tenth of the vector it describes. What costs you is that pgvector's HNSW index keeps its own copy of every vector, so the index is roughly another 61 GB on top of the table — and it is the index that has to stay resident for the latency figures above to hold. Once the working set stops fitting in RAM and index pages start being read back from disk, ANN p95 degrades by an order of magnitude, and it degrades unpredictably rather than gracefully.
That single number reshapes the decision. At 500,000 vectors you are at roughly 3 GB of vectors and a similar amount of index; that lives comfortably inside a Postgres instance you probably already run. At 50 million you are past 300 GB before the index, and now you are talking about either a dedicated store with sharding and quantisation built in, or a Postgres deployment sized and operated like a specialist system, which removes most of the reason to have chosen Postgres.
Quantisation changes the arithmetic, and both camps now support it. Halving to float16 (pgvector's halfvec, or scalar quantisation in Qdrant, Milvus and Weaviate) cuts memory in half for a recall loss most teams cannot detect in their own evaluation set — verify that on yours rather than taking a figure from anyone's blog post, mine included. Binary quantisation cuts it by 32× and is genuinely usable when you rerank the top 100–200 binary candidates against full-precision vectors, which is exactly what a two-stage retrieval pipeline does anyway. The dedicated stores make this a configuration flag; in Postgres it is a schema and query-design decision you own. Neither is wrong, but one of them is more work, and the work is yours.
Where pgvector wins, and it wins often
For a large share of production RAG systems, pgvector is the right answer, and it is the right answer for reasons that have nothing to do with vector search.
The corpus is under 5–10 million vectors. The team already operates Postgres and has backups, replication, monitoring and an on-call rotation that understands it. The vectors live next to the rows they describe, so a retrieval query can join on ownership, permissions and freshness in a single transaction rather than reconciling two systems. Nothing about "we deleted that document" becomes an eventual-consistency problem. And the operational surface area is zero additional services, which for a team of six engineers is worth more than any benchmark.

The HNSW implementation is mature. On a warm index of a few million vectors with ef_search in the 40–100 range, single-digit to low-double-digit millisecond p50 is normal. Build times are the honest weakness: an HNSW index over 5 million 1,536-dimension vectors takes tens of minutes to a few hours depending on maintenance_work_mem and parallelism, and rebuilding it is not something you want on a critical path. Plan index builds as a batch operation, not a deploy step.
Where pgvector stops being the answer
There are three situations where I would not start with pgvector, and none of them is "it is slow".
Narrow filters on a large index. This is the failure mode that produces the 200 ms row in the filtering table. HNSW is a graph traversal, and the naive way to combine it with a WHERE clause is to traverse the graph for the nearest neighbours and then discard the ones that fail the filter. If the filter keeps 30 percent of rows, that works. If the filter keeps 0.1 percent of rows, which is what a per-tenant filter looks like on a shared index, you retrieve your ef_search candidates, throw almost all of them away, and return three results when the user asked for twenty. And because the planner commits up front, from its selectivity estimate, a filter that narrow will often make it skip the vector index altogether and run an exact scan: perfect recall, and on a large table, a latency you cannot ship. pgvector 0.8 added iterative index scans that keep walking the graph until enough results pass the filter, and that helps materially, but it trades latency for recall on exactly the queries that were already your slowest.
The dedicated stores went at this problem earlier. Qdrant, Weaviate and Milvus each ship some form of filterable HNSW that evaluates the payload filter during traversal rather than after it, backed by payload indexes, so that a filter the store knows is highly selective flips the query into a brute-force scan over the small matching set, which at 0.1 percent of 20 million rows is 20,000 vectors and takes a few milliseconds. The specifics move between releases; check the version you will actually run rather than the blog post announcing the feature. This is the single strongest technical argument for a dedicated store, and it is the one the benchmarks do not show, because benchmarks run unfiltered queries.
Multi-tenancy at real scale. If you serve 27,000 institutions and every retrieval must be scoped to one of them, you have the narrow-filter problem 27,000 times over, plus an isolation requirement that a filter clause alone does not satisfy to an auditor. In Postgres the honest answer is partitioning by tenant, or a per-tenant partial index for the largest tenants, both of which work and both of which you now maintain. The dedicated stores treat multi-tenancy as a first-class concern: tenant-aware payload partitioning in Qdrant, per-tenant shards in Weaviate, partition keys in Milvus. The point is not that Postgres cannot do it. It is that the dedicated store has already made the design decisions you are about to make, and made them for this workload.
Corpus growth you cannot bound. A single Postgres node has a memory ceiling, and past it you are sharding Postgres by hand or moving to Citus, and either way you have become a distributed-systems team to keep a vector index running. The dedicated stores shard horizontally as a matter of course. If your roadmap has the corpus growing 10× in eighteen months, the migration you are avoiding today is the migration you will do under load later.
The decision, as I actually make it
Write down the corpus size at launch and at eighteen months. Write down the most selective filter a real query will apply, as a fraction of the index. Write down the p95 you have promised.
If the eighteen-month corpus is under 10 million vectors and the most selective filter keeps more than a few percent of rows, use pgvector, and spend the engineering time you saved on retrieval evaluation, which is where quality is actually decided. Those thresholds assume full-precision vectors; quantise and they move, which is the argument for doing the arithmetic rather than memorising a number.
If the corpus is heading past 20 million, or the selective filter is a per-tenant scope on a shared index, or you need multi-tenancy that an auditor will sign off, start on a dedicated store. Choose between them on operational grounds: whether you want it managed or self-hosted, what your team already knows, and whether the pricing model penalises your read pattern. The pure-performance differences between the serious options are smaller than the differences in how they behave when a node fails.
Between those two cases, and there is a real middle, measure. Load a representative slice of your data, run your ten ugliest real queries with their real filters, and look at the p95, not the p50. This takes a day and settles arguments that otherwise run for weeks.
What I got wrong, for balance
I have under-sized pgvector deployments twice by treating the memory arithmetic as approximate; it is not, and the cliff when the index leaves memory is sharp. I have also over-engineered once by putting a dedicated store in front of a 400,000-vector corpus that never grew, and spent a year operating a service that a Postgres extension would have handled with a fraction of the attention. The second mistake was cheaper than the first, but it was still a mistake, and it came from choosing based on where I thought the product would be rather than where the numbers said it was.
The store is rarely the bottleneck. The decision matters because the wrong one makes the bottleneck yours to operate.
Illustrations generated with AI.
