Why “similar” is not enough

A customer tells a supplier: “I need a bearing for a pump, but the marking on the old one is almost gone.” Another customer provides a precise part number. These are two different workflows. In the first, the seller must find candidates by purpose, dimensions, and technical properties. In the second, the system must verify the identifier before suggesting “something like it.” A search service that treats both requests alike can look impressive in a demo and still propose an incompatible part in a live order.

Dense semantic vectors can connect different ways of expressing the same need. Yet alphanumeric codes, units, and revision suffixes may not be represented reliably by them. Lexical search returns documents containing matching terms. An exact filter on the SKU field answers a stricter question: does this particular code exist? Qdrant’s documentation separates these functions: keyword fields are for exact values, text search handles words and phrases, and hybrid queries combine semantic and lexical results. Those are tool capabilities, not a quality guarantee for a given catalog.

There is a gentle workplace joke with a real operational point: a diligent robot may bring an entire tray of “almost identical” flanges. A person still signs off on the invoice. The interface therefore needs to show which result is an exact match, which is a proposed substitute, and which is merely a lead requiring more information.

Three branches of one search flow

The practical architecture starts with parsing the request, not selecting a model. If it contains a SKU, barcode, or internal supplier code, normalize it under agreed rules for case, spaces, permitted separators, and known prefixes. Do not remove every hyphen or leading zero automatically: for some manufacturers they distinguish products. Keep the original input alongside the normalized value for auditability.

Then route the request as follows:

  • First, seek an exact match in an indexed `sku` field or a code-mapping table. Return the product card, catalog version, and availability status. A semantic model must not silently replace a found code with another one.
  • If no exact match exists, use lexical retrieval for markings in descriptions and documents, and semantic retrieval for intended use and synonyms. Merge their results only as candidates clearly marked “verification required.”
  • For descriptive requests without a code, run semantic and lexical retrieval together, applying category and access filters. Then check mandatory properties and compatibility with independent technical rules.

For a code such as `AB-1200`, do not assume a text analyzer will preserve the hyphen exactly as the business requires. Keep critical identifiers in a separate keyword field or master table instead of relying solely on BM25. Qdrant explicitly distinguishes untokenized keyword strings from tokenized text fields. Full-text search is not the same thing as exact SKU equality.

Where the hybrid layer belongs

In the hybrid branch, one product card can have a dense vector for its semantic description and a sparse representation for words. A query runs two candidate searches and fuses their rankings. Qdrant documents this mechanism through its Query API; reciprocal rank fusion, or RRF, is one available method. RRF uses a document’s rank in the two lists instead of directly adding incomparable dense-model and BM25 scores. It is a sensible starting point, but its final order must be tested on your own queries.

There is a crucial limit: if the required item appears in neither candidate list, fusion cannot invent it. If an exact SKU is already known, asking RRF to guess it among similar products is counterproductive. Hybrid retrieval helps unpack an unclear natural-language need, such as “a high-temperature seal for a food-processing pump.” It does not replace exact lookup or engineering constraints. A popular product is not automatically a compatible one.

For an on-premises setup, place the components near the catalog: the product database or read-only replica, an embedding service, a search index, and the application used by staff. Some Qdrant examples use cloud inference; those examples are not a ready-made prescription for a closed environment. The documentation also describes client-side vector generation, for example with FastEmbed. Choose a particular model only after testing Russian, domain terminology, size, and license—not from a public leaderboard alone.

Data preparation beats clever ranking

First define a canonical product card. At minimum it should contain a stable internal ID, manufacturer and supplier codes, source version and date, name, category, measurable properties with units, compatibility, lifecycle status, and a link to the authoritative document. Live prices, stock, and sales terms are better fetched from the current ERP when displaying a result than baked into an infrequently refreshed vector. Dealer or branch access rights must be enforced before retrieval, not hidden inside a prompt.

A separate alias table helps when an item has old, new, and supplier-specific codes. But “alternative” must not automatically become “same product.” Revision, voltage, material, tolerance, diameter, certification, and operating conditions may determine compatibility. Check these fields with rules or a responsible specialist. A model can explain what further details to ask for; it must not claim compatibility without evidence.

A Russian-language catalog needs explicit tests for morphology and tokenization. Qdrant warns that default BM25 settings are oriented to English. Other languages need appropriate processing applied consistently at indexing and query time. Test mixed Russian-English descriptions, abbreviations, and factory codes separately. Poor normalization can damage exact retrieval faster than a new ranking algorithm can repair it.

A pilot that can disprove the idea

Start with 50–100 anonymized real queries if that volume is available. Split them into at least four groups: exact code, mistyped code, description without a code, and request for a substitute. For each, record the correct item or the honest outcome “needs clarification,” plus mandatory compatibility attributes. This is an evaluation set, not training data that promises a result. The catalog owner and technical specialist should label it and adjudicate disputed cases.

On the same set, compare exact lookup, lexical retrieval, semantic retrieval, and hybrid retrieval. Measure more than the rate of relevant products near the top. Track false “exact” matches, incompatible proposals, correctly abstained queries, 95th-percentile latency, and staff time needed to confirm an item. Reducing wrong SKU recommendations matters more than winning an attractive average metric on descriptive questions. If hybrid retrieval does not improve the relevant query classes, the extra index is unnecessary.

Calculate economics as a model with your own assumptions. Cost per accepted result equals the cost of preparing and updating cards, infrastructure, and review time, divided by confirmed useful answers. Separately estimate the cost of a wrong shipment, return, and delayed order. We offer no universal currency figure or claimed speed-up: the outcome depends on assortment, master-data quality, and approval process. A pilot may reveal that cleaning codes delivers more value than adding another model.

A low-cost next step

Pick one product family with frequent questions and collect benchmark queries. Implement exact SKU lookup first, then both retrieval branches over the same product cards. Do not enable automatic order creation. Show staff the origin of each result, its source document, compatibility parameters, and the reason for uncertainty. After two weeks, compare errors and confirmation time with the existing process. Only then decide whether a hybrid layer belongs in the production system.

The architectural takeaway is simple: a local model helps when a buyer describes a task in ordinary language. When the buyer names an exact code, strict equality and an up-to-date catalog matter more. The robot may bring options; the specification and responsibility for the order should remain with people.