Why replacing the model file is not enough
A local RAG assistant is already answering employees' questions about procedures, contracts, and the knowledge base. The team finds a new embedding model that understands Russian queries better or runs faster on existing hardware. It is tempting to change the model in the configuration and keep the old index. That is a poor migration. Document vectors and query vectors must be produced by a compatible pair; changing the model means re-embedding the existing documents. The vector dimension may change as well, and the new collection must match it.
The business risk is more than a technical error. The assistant may remain available while becoming worse at finding the right policy, missing a current document, or citing a less relevant one. Users will notice before CPU monitoring does. “Zero downtime” therefore means both API availability and acceptable answer quality during the switch.
Qdrant's official guide describes two routes: build a new collection alongside the old one and redirect queries, or—given an appropriate version and a collection originally created with named vectors—add a second vector representation inside the collection. Weaviate's independent guide likewise recommends comparing quality on the same sample before switching and using an alias for a reversible cutover. Exact commands vary by store, but the management principle is the same: measure and build first, switch later.
Prepare before migrating
You need the original text from which vectors were generated. Old numbers alone cannot produce embeddings from a new model. The text may live in the document system or the index payload; it is better to treat the document system as the source of truth and keep stable `document_id`, `chunk_id`, version, and source references in the index. If chunking rules change at the same time, comparing models becomes harder: both the model and the corpus have changed. For an initial migration, keep the documents and chunks identical.
Record the following before starting:
- exact old and new model versions, tokenizers, and encoding settings, plus vector dimensions and search distance;
- document and chunk counts, missing items, deleted versions, and access restrictions;
- 50–100 real but anonymized questions, each with the correct document and expected answer;
- the share of questions whose correct document appears among the first search results, final answer quality, search latency, and cost per accepted answer;
- the maximum acceptable rollback time and the person authorized to switch search traffic.
The suggested question count is a practical size for a small pilot, not a statistical guarantee. Include rare product codes, similar names, Russian and English documents, recent changes, and unanswerable questions. For critical processes, test permissions too: the new index must not widen document visibility.
The straightforward route: two collections, one search address
The clearest small-business design is a blue-green index. The old collection serves users. A new collection is created separately with dimensions and distance metric appropriate to the new model. A background worker reads the source chunks, computes new vectors, and writes the same identifiers and access metadata to the new collection. Ordinary queries continue to use the old collection while the new one fills.
Changing documents cannot be ignored during a long rebuild. From the start of migration, new records and edits must reach both indexes or a durable queue that the new version catches up with before cutover. Qdrant explicitly warns that the simple dual-write pattern in its example handles whole-point upserts, but deletes and partial updates require extra handling or a temporary pause. Otherwise a background pass can reintroduce a document that was already deleted. For contracts or personal data, that is more than a relevance problem.
After backfilling, compare expected object counts and checksums; spot-check versions and permissions; run the same evaluation questions against both indexes. Then switch the production alias or search configuration to the new collection together with the model that encodes queries. Qdrant's alias operations are atomic, so a query should not land on a half-switched collection address. But the alias does not switch the query encoder itself. The application must change the “query encoder + index” pair coherently, for example through versioned configuration or a route that selects both as one release.
Do not delete the old collection immediately. It is the rollback path while the new search is observed on real but safe traffic. Rollback means restoring the old “query model + collection” pair, not just changing an alias. Then stop dual writes and remove old data only under the agreed retention policy and after checking the backup.
When named vectors fit
Qdrant documents another route for collections with named vectors on version 1.18 or newer. Add a new named vector field; new records receive both representations; existing points are re-embedded in the background; then queries switch their `using` parameter and query model together. Existing IDs and payload metadata stay in place, and the old vector can be retained temporarily for rollback. This duplicates less text metadata, but the transition still needs compute capacity and space for two vector sets.
It is not a universal button. If the collection was created without named vectors or the server version is below the one in the guide, use parallel collections. Do not assume one product's behavior applies to another: Weaviate's guide, for example, describes its own limitations around adding and removing vectors. Check the exact deployed version, schema, and operating mode rather than a generic slide-deck recommendation.
Quality gates matter more than an elegant cutover
An “API returned 200” test says nothing about whether knowledge became easier to find. Compare old and new pairs on the same sample: correct document among top results, completeness of citations, share of human-accepted answers, errors on item numbers or customer names, and 95th-percentile latency. Do not change the embedding model, reranker, and answer template in one experiment; the cause of any improvement or regression would be unclear.
Test three awkward events separately. First, a document changes during the background pass: the new collection must end up with the latest version. Second, a document is deleted or access is revoked: the new version must not return it. Third, the re-embedding worker crashes halfway through: resuming must not create duplicates or overwrite fresher data. This calls for a checkpoint, idempotent processing, and a queue-lag check before enabling the new search.
The business can use a simple launch rule: employees get the new model only if it does not degrade critical answers on an agreed test set and stays within the latency limit. Better average retrieval does not justify losing the right answer to a rare but costly question. The process owner and IT should decide together, not an automated benchmark alone.
Cost and the first step
The migration budget includes re-embedding the whole corpus, temporary storage for a second collection or vector, dual processing of new documents, evaluation, and rollback reserve. A useful model calculation is chunk count × measured encoding time per chunk on your own hardware, plus peak space for both indexes and review labor. Do not put someone else's GPU speed into the budget: document length, batch size, CPU/GPU, and competing workload change the result materially.
The first step is not to rebuild the whole knowledge base. Pick one department, a stable document sample, and 50–100 questions. Build a second small collection, measure search quality and cost, then rehearse cutover and rollback on a test address. If quality is no better or the transition cost is not justified, keep the old model. If it is better, you have a verified plan to migrate without shutting down the assistant or silently losing documents.
