A better embedding model is not an upgrade. It is a migration, and the corpus you have already embedded is the legacy system.
There is a great deal published on choosing an embedding model and almost none on replacing one, which is why teams that read the benchmarks carefully still ship the change badly. The failure is not dramatic. Nothing errors, latency does not move, retrieval quality quietly gets worse in a way that is hard to attribute, and someone eventually concludes the new model was overrated.
Here is the plan I would defend in a design review.
The constraint that decides everything else
Two embedding models do not share a vector space. Not approximately, not after normalisation, not if the dimensions happen to match.
This has one consequence that determines the entire design: you cannot re-embed in place, batch by batch, against a live index. A query encoded by the new model and compared against an index that is 60 per cent new vectors and 40 per cent old ones does not return slightly degraded results. The similarity scores against the old-model vectors are arbitrary, so the ranking between the two populations is arbitrary. Depending on the distance metric and how the two models happen to scale, you will get a systematic bias toward one population or the other, and which one is not something you can reason about in advance.
Nothing in your monitoring notices. No exception, no latency change, no drop in result count, because a retrieval system returning the wrong ten plausible documents looks exactly like one returning the right ten.
So the new vectors go somewhere else. Always. The work is a migration between two indexes that briefly coexist, and the engineering problem is keeping them consistent and choosing the moment to switch.
Five things to check before you spend anything on tokens
Most of the expensive surprises in this work are discoverable in an afternoon.
Dimension. The visible change, and the cheap one. It forces a new index and new column definitions, which you needed anyway. Worth noting only because it makes the change feel structural when the structural part is invisible.
Normalisation and metric. Some models emit L2-normalised vectors and are trained for cosine similarity; others are not normalised and expect inner product. If your index is configured for inner product and the new model is not normalised the way the old one was, rankings change for reasons that have nothing to do with model quality. Check what the model card says the metric should be, and check what your index is actually configured with. These disagree more often than they should.
Asymmetric prefixes. Several strong retrieval models are trained asymmetrically: documents are embedded with one instruction prefix and queries with another. Omit the query prefix and recall degrades substantially while everything still runs. This is the single most common way I see a migration misread as a model regression — the model is fine, the query path is half-configured, and the evaluation numbers say "the new model is worse".
Tokenizer and truncation. Your chunks were sized in the old model's tokens. The new model has a different tokenizer, and a 512-token chunk under the old one may be 600 under the new one. Most SDKs truncate to the input limit by default and do not raise. The check is one line: run your chunk length distribution through the new tokenizer and look at the tail before you embed ten million of them, not after.
Input limit and batch shape. Maximum input tokens and maximum batch size both change your throughput, and throughput decides your schedule.
Sizing the job honestly
Take a corpus of 10 million chunks averaging 350 tokens. That is 3.5 billion tokens to embed.
Money is not the constraint. At hosted embedding prices in the band of roughly US$0.02 to US$0.13 per million tokens, the full re-embed costs between about US$70 and US$455 — a rounding error against the engineering time. Worth saying plainly, because teams delay this work on a cost instinct that does not survive the arithmetic.
Throughput is the constraint. If your account is allowed one million tokens per minute, 3.5 billion tokens is 3,500 minutes, or just under 60 hours of continuous embedding. At five million tokens per minute it is about 12 hours. That number decides whether this is an overnight job or a week-long background process, and every other decision (how long you dual-write, how long you carry two indexes, how much drift the backfill reconciles) falls out of it.
Storage peaks higher than people plan for. Ten million 1,024-dimension float32 vectors are about 41 GB of raw vector data. In pgvector, an HNSW index holds its own full copy of every vector, so the index is roughly another 41 GB plus the neighbour lists, and at m = 16 those lists are on the order of 128 bytes per vector — about 1.3 GB, or a few per cent, not the dominant term. Call the new index around 83 GB all in. If the old model was 768 dimensions, the old side is about 63 GB on the same accounting. You need both at once: a peak near 146 GB to end up at 83 GB. Provision for the peak, not the destination.
Index build is its own line item. The HNSW build is not free and it is not linear. Build it on a one per cent sample, time it, and extrapolate with a little pessimism. It also cannot be fully overlapped with the re-embed: you can pipeline the build as batches land, but the tail of the build cannot start until the last batch does, so your critical path is the re-embed plus the build tail, not the larger of the two.
The mechanism: dual-write, watermark, backfill
The shape is the same as any online data migration, and the ordering matters more than the components.
- Create the new index empty. New table or collection, new dimension, metric set to what the new model expects.
- Turn on dual-write first. From this moment, every create, update and delete that touches the corpus writes to both indexes, embedding with the appropriate model for each. This must be live before the next step.
- Take the snapshot watermark. Record the point in the source of truth from which the backfill will read. Because dual-write was already on, anything that changes after this instant is covered; anything before it is covered by the backfill. Take the watermark first and you have a gap you cannot see and cannot recover without a full recompare.
- Backfill from the source of truth, not from the old index. This is the part that looks like a needless extra hop and is not. Which brings us to the trap.
The delete trap. A document is deleted during the backfill. Dual-write dutifully deletes it from both indexes, except it was never in the new index yet, so that delete is a no-op. Then the backfill reaches its batch and happily inserts it. You now have a document in your retrieval index that does not exist in your product, and you will find out when it is quoted back to a user.

Two clean fixes: the backfill reads current state from the source of truth and re-checks existence at write time, or you keep a tombstone set of everything deleted after the watermark and consult it before every insert. The first is simpler, and is why step four says to backfill from the source of truth. Updates are the same shape in a milder form; both are resolved by making the backfill write conditional on the record still existing and not having been written more recently by the live path.
Proving the new model is actually better
You should not cut over on a benchmark. Benchmarks are chosen by the people selling the model, and your corpus is not their corpus.
Freeze the queries — a few hundred real ones drawn from production traffic across your actual query mix, not the ten anyone can remember. What you must not freeze is the judged document set.
If your relevance labels were mined from old-system behaviour — clicks, thumbs, which chunk the answer cited — then your labels only cover documents the old model was capable of retrieving. Score the new model against that set and it is penalised for every good document the old system never surfaced, because an unjudged document scores as irrelevant. That is pooling bias, it has been understood in information retrieval for decades, and almost nobody applies the correction when swapping an embedding model.
The correction is cheap: for each frozen query, pool the top-k candidates from both systems, judge the union blind to which system produced each, and only then compute metrics. A few hundred queries at k=10 is a few thousand judgements — an afternoon with a rubric or a well-anchored judge model. The method for building and running that set is the one I set out in the agent-evaluation piece, so I will not restate it.
Report recall at your real k — the k your prompt assembly actually uses, not k=100 — and report it per query class. An aggregate that improves while your highest-value class regresses is the outcome this whole exercise exists to catch.
Cutover, and why it is a feature flag
Once the backfill is complete and the offline numbers hold, run shadow reads: every live query is embedded with both models and run against both indexes, the old result is served, the pair is logged. Watch one number above all — overlap between the two result sets at your real k, sliced by query class. Near-total overlap means nothing changes for those users. A class where overlap collapses is where the risk lives, and it deserves a human eye on twenty examples before anyone flips anything.
Then cut over as a read-path change: one flag that switches the query encoder and the index it queries, together and atomically. Not a deploy. A flag, ramped by traffic slice, because the two must never be out of step for a single request.
The rollback nobody actually has
Most teams stop dual-writing at cutover. It is the natural instinct — the migration is over, why pay for two writes.
The old index then decays from that moment, and the rollback you believe you have is a rollback to a stale corpus, which by day three is its own incident. Decide the window explicitly, keep dual-writing to the old index for its whole duration, and diarise the date you turn it off. A fortnight is usually generous, and two writes for fourteen days is the cheapest insurance in the project.
What I got wrong
I have run a re-embed in place, batch by batch, against a live index, on the reasoning that each converted document was strictly better than before. Quality degraded for the converted portion specifically, no dashboard moved, and it took far too long to find, because every instinct says to look at the model and the answer was in the mixture.
And I have run a backfill from a snapshot while deletions continued against the source of truth, and had deleted documents reappear in retrieval. The backfill was correct. The snapshot was correct. Neither of them knew about the other.
The short version
Two models do not share a space, so the new vectors go in a new index and the query encoder switches with it, atomically. Check normalisation, prefixes and tokenizer before you spend anything, because those three account for most migrations that get misread as the model being worse. Size the job by throughput and peak storage rather than by cost. Dual-write before you take the watermark, backfill from the source of truth rather than from the old index, and handle deletes deliberately. Freeze the queries, pool the candidates, judge blind. Cut over behind a flag after shadow reads, and keep the old index writable for a defined rollback window.
None of this is difficult. All of it is ordering, and the ordering is the migration.
Illustrations generated with AI.
