Most vector-database decisions I have reviewed were made in the wrong order. The team read a benchmark, picked a store, built the retrieval pipeline on top of it, and then discovered what their latency budget actually was when the first users complained. By that point the store was load-bearing and the conversation had become about tuning rather than choosing.

The order that works is the reverse. Write the latency budget first. Size the index with arithmetic, not hope. Then let those two numbers pick the store, because in most cases they pick it for you, and the benchmark you were about to read becomes irrelevant.

Step one: the latency budget, written down

A retrieval-augmented request has five or six stages, and only one of them is the vector search. For a chat-style product with a 1,200 ms p95 target to first token, a realistic budget looks like this:

StageTypical p50Typical p95Notes
Query embedding (hosted API)40 ms120 msNetwork round trip dominates; self-hosted small models get this to 5–15 ms
ANN search (top-k, k=20–50)5–20 ms30–60 msWarm index, in memory
Metadata filtering0–5 ms5–200 msNot a separate stage — this is what the filter adds to the search. The wide range is the whole article
Reranking (cross-encoder, 20–50 candidates)60 ms180 msOften the largest non-LLM slice
Prompt assembly2 ms5 msNegligible unless you are doing something odd
LLM time-to-first-token300 ms700 msThe budget you do not control

Those figures are a worked example built from public sizing arithmetic, not a trace from a system I run. Copy the shape, not the numbers. And note that the p95 column sums to more than 1,200 ms. That is the correct and uncomfortable result: stage p95s do not co-occur, so the sum is a worst case rather than the end-to-end p95, but it tells you there is no slack for a stage that misbehaves.

Two things fall out of writing this table. First, the ANN search is around 5 percent of the budget on a good day. A store that is twice as fast at pure vector search saves you single-digit milliseconds at p50 and perhaps thirty at p95. That is not a reason to choose anything. Second, the line with the widest variance is metadata filtering, and it is the one most teams do not think about until the p95 goes wrong. Filtering behaviour, not raw search speed, is the real axis of the pgvector-versus-dedicated decision.

Bar chart of the article's worked example, p95 latency by stage for a 1,200 ms target to first token: query embedding 120 ms, ANN search 30 to 60 ms, metadata filtering 5 to 200 ms, reranking 180 ms, prompt assembly 5 ms and LLM time to first token 700 ms. The search is a small slice; the filter has the widest range. p95 by stage, for a 1,200 ms target Query embedding, hosted API 120 ms ANN search, warm index 30–60 ms Metadata filtering, added to the search 5–200 ms Reranking, cross-encoder 180 ms Prompt assembly 5 ms LLM time to first token 700 ms Two-tone bars show a range, low to high. Worked example from the text, not a trace.
On a good day the vector search is around 5 percent of the budget; the filter is the line that can add 200 ms. Worked example from the text.

Step two: size the index honestly

The arithmetic is simple and almost nobody does it before choosing.

A 1,536-dimension float32 embedding is 6,144 bytes. Ten million of them is 61 GB of raw vectors. The neighbour lists are not where the memory goes, which is the part people get wrong: at a typical m of 16, each node carries a few dozen four-byte links, well under a tenth of the vector it describes. What costs you is that pgvector's HNSW index keeps its own copy of every vector, so the index is roughly another 61 GB on top of the table — and it is the index that has to stay resident for the latency figures above to hold. Once the working set stops fitting in RAM and index pages start being read back from disk, ANN p95 degrades by an order of magnitude, and it degrades unpredictably rather than gracefully.

That single number reshapes the decision. At 500,000 vectors you are at roughly 3 GB of vectors and a similar amount of index; that lives comfortably inside a Postgres instance you probably already run. At 50 million you are past 300 GB before the index, and now you are talking about either a dedicated store with sharding and quantisation built in, or a Postgres deployment sized and operated like a specialist system, which removes most of the reason to have chosen Postgres.

Bar chart of the article's sizing arithmetic for 1,536-dimension float32 vectors at 6,144 bytes each: 500,000 vectors take about 3 GB plus a similar index; 10 million take 61 GB plus about 61 GB for pgvector's HNSW index, which keeps its own copy and must stay in RAM; 50 million pass 300 GB before the index. pgvector memory at 1,536 dimensions float32: 1,536 × 4 = 6,144 bytes per vector 500,000 vectors ≈ 3 GB plus a similar amount of index 10 million vectors 61 GB + ≈ 61 GB table index: its own copy, must stay in RAM 50 million vectors past 300 GB before the index Out of RAM, ANN p95 degrades by an order of magnitude, and unpredictably. Worked example from the text.
Do the arithmetic before choosing: pgvector's HNSW index holds its own copy of every vector, and it has to stay in memory. Worked example from the text.

Quantisation changes the arithmetic, and both camps now support it. Halving to float16 (pgvector's halfvec, or scalar quantisation in Qdrant, Milvus and Weaviate) cuts memory in half for a recall loss most teams cannot detect in their own evaluation set — verify that on yours rather than taking a figure from anyone's blog post, mine included. Binary quantisation cuts it by 32× and is genuinely usable when you rerank the top 100–200 binary candidates against full-precision vectors, which is exactly what a two-stage retrieval pipeline does anyway. The dedicated stores make this a configuration flag; in Postgres it is a schema and query-design decision you own. Neither is wrong, but one of them is more work, and the work is yours.

Where pgvector wins, and it wins often

For a large share of production RAG systems, pgvector is the right answer, and it is the right answer for reasons that have nothing to do with vector search.

The corpus is under 5–10 million vectors. The team already operates Postgres and has backups, replication, monitoring and an on-call rotation that understands it. The vectors live next to the rows they describe, so a retrieval query can join on ownership, permissions and freshness in a single transaction rather than reconciling two systems. Nothing about "we deleted that document" becomes an eventual-consistency problem. And the operational surface area is zero additional services, which for a team of six engineers is worth more than any benchmark.

A drawer of pale folders, each with a small teal clip attached; one orange folder is lifted out and its clip comes with it, standing for vectors kept beside the rows they describe, so that deleting a document never becomes an eventual-consistency problem.
pgvector wins for reasons that have nothing to do with vector search. The vectors live next to the rows they describe, and a deleted document takes its vector with it.

The HNSW implementation is mature. On a warm index of a few million vectors with ef_search in the 40–100 range, single-digit to low-double-digit millisecond p50 is normal. Build times are the honest weakness: an HNSW index over 5 million 1,536-dimension vectors takes tens of minutes to a few hours depending on maintenance_work_mem and parallelism, and rebuilding it is not something you want on a critical path. Plan index builds as a batch operation, not a deploy step.

Where pgvector stops being the answer

There are three situations where I would not start with pgvector, and none of them is "it is slow".

Narrow filters on a large index. This is the failure mode that produces the 200 ms row in the filtering table. HNSW is a graph traversal, and the naive way to combine it with a WHERE clause is to traverse the graph for the nearest neighbours and then discard the ones that fail the filter. If the filter keeps 30 percent of rows, that works. If the filter keeps 0.1 percent of rows, which is what a per-tenant filter looks like on a shared index, you retrieve your ef_search candidates, throw almost all of them away, and return three results when the user asked for twenty. And because the planner commits up front, from its selectivity estimate, a filter that narrow will often make it skip the vector index altogether and run an exact scan: perfect recall, and on a large table, a latency you cannot ship. pgvector 0.8 added iterative index scans that keep walking the graph until enough results pass the filter, and that helps materially, but it trades latency for recall on exactly the queries that were already your slowest.

The dedicated stores went at this problem earlier. Qdrant, Weaviate and Milvus each ship some form of filterable HNSW that evaluates the payload filter during traversal rather than after it, backed by payload indexes, so that a filter the store knows is highly selective flips the query into a brute-force scan over the small matching set, which at 0.1 percent of 20 million rows is 20,000 vectors and takes a few milliseconds. The specifics move between releases; check the version you will actually run rather than the blog post announcing the feature. This is the single strongest technical argument for a dedicated store, and it is the one the benchmarks do not show, because benchmarks run unfiltered queries.

Diagram of a query whose filter keeps 0.1 percent of rows. Filtering after the HNSW search discards almost every candidate and returns 3 of the 20 results asked for; exact and iterative scans buy recall back with latency. Filtering during the search, as dedicated stores do, brute-forces the 20,000 matching vectors of 20 million in a few milliseconds. A filter that keeps 0.1 percent of rows Filter after the search HNSW, then WHERE 3 results of the 20 asked for Exact scans, or pgvector 0.8 iterative scans, buy the recall back with latency. Filter during the search dedicated stores a selective filter flips to a brute-force scan 0.1% of 20 million rows is 20,000 vectors a few milliseconds
This is the strongest technical argument for a dedicated store, and the one benchmarks miss, because they run unfiltered queries.

Multi-tenancy at real scale. If you serve 27,000 institutions and every retrieval must be scoped to one of them, you have the narrow-filter problem 27,000 times over, plus an isolation requirement that a filter clause alone does not satisfy to an auditor. In Postgres the honest answer is partitioning by tenant, or a per-tenant partial index for the largest tenants, both of which work and both of which you now maintain. The dedicated stores treat multi-tenancy as a first-class concern: tenant-aware payload partitioning in Qdrant, per-tenant shards in Weaviate, partition keys in Milvus. The point is not that Postgres cannot do it. It is that the dedicated store has already made the design decisions you are about to make, and made them for this workload.

Corpus growth you cannot bound. A single Postgres node has a memory ceiling, and past it you are sharding Postgres by hand or moving to Citus, and either way you have become a distributed-systems team to keep a vector index running. The dedicated stores shard horizontally as a matter of course. If your roadmap has the corpus growing 10× in eighteen months, the migration you are avoiding today is the migration you will do under load later.

The decision, as I actually make it

Write down the corpus size at launch and at eighteen months. Write down the most selective filter a real query will apply, as a fraction of the index. Write down the p95 you have promised.

If the eighteen-month corpus is under 10 million vectors and the most selective filter keeps more than a few percent of rows, use pgvector, and spend the engineering time you saved on retrieval evaluation, which is where quality is actually decided. Those thresholds assume full-precision vectors; quantise and they move, which is the argument for doing the arithmetic rather than memorising a number.

If the corpus is heading past 20 million, or the selective filter is a per-tenant scope on a shared index, or you need multi-tenancy that an auditor will sign off, start on a dedicated store. Choose between them on operational grounds: whether you want it managed or self-hosted, what your team already knows, and whether the pricing model penalises your read pattern. The pure-performance differences between the serious options are smaller than the differences in how they behave when a node fails.

Between those two cases, and there is a real middle, measure. Load a representative slice of your data, run your ten ugliest real queries with their real filters, and look at the p95, not the p50. This takes a day and settles arguments that otherwise run for weeks.

Grid of the article's decision rule, with corpus size at eighteen months across and the most selective filter down. Under 10 million vectors with a filter that keeps more than a few percent of rows: pgvector. Past 20 million, or a per-tenant scope on a shared index: a dedicated store. The highlighted cells in between: measure. Where corpus size and filter point Most selective filter Corpus at eighteen months under 10M 10M to 20M past 20M keeps more than a few percent keeps a few percent or less per-tenant scope, shared index pgvector measure dedicated measure measure dedicated dedicated dedicated dedicated Measure: ten ugliest real queries, at p95. Audited tenant isolation: dedicated, too. Schematic; thresholds assume full precision.
Under 10 million vectors with broad filters, pgvector; past 20 million or a per-tenant scope, a dedicated store. In between, a day spent on your ten ugliest real queries at p95 settles it.

What I got wrong, for balance

I have under-sized pgvector deployments twice by treating the memory arithmetic as approximate; it is not, and the cliff when the index leaves memory is sharp. I have also over-engineered once by putting a dedicated store in front of a 400,000-vector corpus that never grew, and spent a year operating a service that a Postgres extension would have handled with a fraction of the attention. The second mistake was cheaper than the first, but it was still a mistake, and it came from choosing based on where I thought the product would be rather than where the numbers said it was.

The store is rarely the bottleneck. The decision matters because the wrong one makes the bottleneck yours to operate.

Illustrations generated with AI.