The business problem

A small company sells imported equipment. Engineers answer customers in Russian, while some manuals, service bulletins and compatibility tables arrive in English. ‘Translate everything and load it into RAG’ sounds natural. But translating a knowledge base and translating each question are different cost items, while direct multilingual search is not a free guarantee of quality. The real question is where the correct passage lives and what it costs to deliver it to an employee with a link to the original.

RAG is not an independent technical expert in this scenario. It retrieves relevant passages and passes them to a model to draft an answer; a person remains responsible for the decision and for sending it. If the index supplies the wrong document, eloquent wording cannot repair the mistake. Part numbers, operating warnings, units of measure and manual revisions are particularly sensitive.

Three routes worth comparing

**1. Translate documents before indexing.** Originals remain the source of truth; approved translations sit alongside them with matching section identifiers. A Russian query searches the Russian copy. The advantage is understandable results for staff. The disadvantages are the initial translation of a large corpus, retranslating changes, storing parallel versions and editorial review of critical passages. Unchecked machine translation can alter a negation or part number.

**2. Translate each query.** Search runs over the original English documents; each Russian question is translated at request time. There is no initial translated copy, and updates can be indexed without a translation step. The price is an extra step on every request, added latency and the risk of mistranslating the user's term. The final answer still needs to be written in Russian and checked against the English source.

**3. Search directly with multilingual embeddings.** Documents and queries are encoded by one compatible model into vectors, while the index keeps passages in their original language. This removes mandatory input translation, but quality must be measured on the company's actual language pairs. The BGE-M3 model card says it supports more than 100 languages and inputs up to 8,192 tokens, as well as dense and sparse retrieval; the model is MIT-licensed. These are tool properties, not a promise of equal accuracy for Russian questions over a particular set of manuals.

In a September experiment on the XRAG dataset, Qdrant reported that multilingual-e5-small placed a relevant same-language passage in the top ten in 61% of cases, but an equally relevant passage in another language in only 8%. This is one dataset and one model, not a forecast for a Russian business. The separate XRAG study also shows that, even after retrieval, a generator can mishandle the response language or reasoning across documents in different languages. Retrieval and answer quality therefore need separate tests.

An illustrative calculation: where the expense appears

Consider a **hypothetical** workload, not a real company's outcome: 20,000 manual passages averaging 1,000 characters each, 500 queries per business day, 22 business days per month, and 10% of the corpus updated monthly. For the calculation, a question averages 100 characters. We deliberately do not insert prices for translation, hardware or labour: those depend on deployment, contract and quality.

  • Translating the whole corpus processes roughly 20 million characters once; translating the 10% update adds 2 million characters per month. Review and storage of parallel versions come on top.
  • Translating queries processes 500 × 22 × 100 = 1.1 million characters per month. Translation happens during the request and contributes to answer latency.
  • Direct multilingual retrieval needs no such translation for retrieval, but it does require encoding 20,000 passages once, encoding each new query, operating an index, and checking cross-lingual retrieval. A Russian answer based on an English passage may still need generation or translation and human control.

The ‘2 million versus 1.1 million’ comparison does not mean an automatic near-halving of cost. Passages and questions differ in complexity, tariff and reuse, while editorial checking can cost more than computation. The calculation simply exposes the workload before product selection. A monetary model must separately include prices per character or token, reviewer time, operation of existing infrastructure, storage and the cost of wrong answers.

It is more useful to calculate **cost per accepted answer** than cost per model call: divide all period costs by the number of answers an employee accepts after checking the source. If a cheap route sends more requests to people or produces incorrect instructions, its cost per usable answer rises. A pilot should also log time to the first usable passage and the share of honest abstentions when no supporting source is found.

A short pilot design

Start with one section of the knowledge base, such as service bulletins for one equipment line. For every passage retain its language, product model, revision number, effective date, original identifier and access level. Remove duplicates and obsolete revisions from active retrieval, but retain them for audit. Restricted documents must not enter a common index without permission filters.

Next, collect 60–100 real anonymised questions. Include Russian questions answered only by English documents, questions with exact part numbers, and questions with no answer in the corpus. Annotate the correct passage and valid document revision manually. Run the same set through all three routes. Measure correct-source recall in the top five results, actual latency, the share of answers with correct citations, human corrections and abstentions. Evaluate cross-lingual retrieval separately: a high overall score can hide failure specifically on Russian–English pairs.

If the source list is small and answers usually lie in documents of the same language as the query, simple retrieval over originals may have the best total cost. If Russian wording is not sufficient to find an exact English part number, add exact-match lookup or a catalogue filter. If cross-lingual retrieval is weak, try translating only the query or a limited set of critical documents—do not translate the whole archive by habit. Additional reranking makes sense only after first-stage search retrieves the correct document at all.

Limits and a decision

A local model makes sense when manuals and questions cannot be sent to an external service, or workload justifies operating an in-house environment. But ‘local’ alone does not solve permissions, bulletin freshness, translation quality or control of generated answers. Measure available memory and utilisation of existing infrastructure on the actual corpus; do not buy a server from a model-card size. First measure speed and quality on the pilot set.

The editorial conclusion for a small business is to spend a week comparing the three routes on a limited corpus rather than mass-translating an archive. Approve quality and response-time thresholds in advance. After the pilot, choose the least expensive route that delivers a verifiable source within the working SLA, while keeping an employee able to stop a wrong answer. The robot intern may handle retrieval; a stack of translated copies reaching the ceiling is a poor KPI.