What the team actually tested
A 2025 study of the EASI-RAG method describes a French environmental analysis laboratory with about 200 employees. Its operators needed answers from operating procedures. The researchers and company tested whether a RAG assistant could find the right passages quickly and answer routine staff questions. The laboratory is unnamed and its documents are not public for confidentiality reasons. This is one documented organisation, not a promise for every laboratory.
The team comprised two users representing a group of 15, an experienced procedure expert, the department head and two developers from the internal IT team. The developers had no previous RAG experience. The corpus was modest: nine files, seven Word and two Excel, totalling 137 pages and about 29,700 tokens. One Word document alone had 95 pages. Some files contained images; some tables were straightforward, while others used merged cells and colour to convey meaning.
The business task was narrow: reduce the time needed to find an applicable rule without giving an operator an instruction that contradicts the procedure. The assistant did not replace laboratory controls, calculate analysis results or change the procedures. Similar tasks for Russian SMEs include service instructions, production routings and internal quality standards, provided the documents are approved and role-based access is enforced.
Why the first version made mistakes
The researchers started with a simple pipeline: file loading, fixed chunks of roughly 1,000 tokens, embedding search, three retrieved passages and answer generation. Embeddings used the compact all-MiniLM-L6-v2 model. Document processing and retrieval ran on an ordinary computer with 8 GB of RAM and no GPU; answer generation used the external gpt-3.5-turbo API. Calling the entire experiment local would therefore be inaccurate.
Users wrote 42 realistic questions plus eight whose answers were absent from the documents. An expert supplied the expected answer and its location in advance. On the first 50-question check, the paper reports 17 correct answers, 19 acceptable partial answers or expressions of uncertainty, and 14 incorrect answers. Average response time was about two seconds. Speed did not solve the real problem: a wrong instruction in an operating procedure matters more than fluent dialogue.
The error review found mundane causes. Once the first Excel sheet was chunked, column headers appeared in the first chunk while later rows lost their context. Another sheet with merged cells, blanks and colour coding did not automatically become intelligible text. Rare technical terms were missed by semantic retrieval. Some chunks broke in the middle of important sentences, and the generator sometimes supplied knowledge that was absent from the procedures.
The fixes were in data and evaluation, not a bigger model
The team prepared the tables: for one Excel file, each row became a separate chunk with column names repeated. They manually converted a second short sheet with colour and marks into textual descriptions. Documents were then split by sections and subsections instead of token count alone. BM25 was added alongside dense retrieval so rare exact terms would not disappear; the number of retrieved passages was later increased. The generator was instructed to use only the supplied evidence.
According to the paper, average response time remained around two seconds after these changes. The authors report 44 correct, seven acceptable and zero incorrect answers. There is an editorial caveat: these categories total 51 although the original test set is described as 50 questions. The paper does not explain that discrepancy, so we do not recalculate a precise final success rate; “zero incorrect” applies only to the described controlled check. The independent ACL paper RAGTruth is a reminder that retrieval may reduce, but does not eliminate, unsupported claims.
The case itself shows the boundary of a closed test. After the tool went live, staff found two new questions that produced incorrect answers and added them to the regression set. Users also asked to see the source passage and document so they could verify answers. An error-report button was added: an expert first helps the user, then checks whether the answer is missing from the corpus or developers need to investigate. That is more useful than claiming mistakes vanished after the pilot.
What we know about cost, and what we do not
Three weeks elapsed from project kickoff to the first production version. The authors estimate about 70 person-hours of team effort, including roughly 50 hours from two developers. The paper mentions modest direct API spending, but that amount belonged to a particular model and price period; it cannot be transferred to another year, provider or local model. Operator time, ongoing maintenance, procedure updates and the cost of a possible mistake require separate accounting.
The study does not provide a reliable like-for-like comparison of search time before and after deployment, nor a validated return-on-investment calculation. The defensible conclusion is not “RAG saved a specific amount”, but “a small team built a testable process and found the main data bottlenecks”. For your own pilot, measure manual search time, the share of correct answers backed by approved sources, dangerous errors, expert maintenance time and infrastructure cost. Track questions for which the system correctly says it does not know.
Adapting the approach to a Russian-controlled environment
Start with approved procedures and someone accountable for keeping them current. Select several dozen real questions, including questions the documentation cannot answer. A process expert should specify the expected answer, document version and acceptable refusal in advance. Define access rights and query logging before inviting employees: search must not expose documents a person is not entitled to read.
If documents cannot be sent to an external API, a local generator and index may be appropriate, but that is a new architecture, not a copy of the French experiment. Replacing the generator can change quality, latency and cost, so run the same questions again. For a small pilot, use already approved storage and an existing search service if their isolation and capacity fit. The word RAG alone does not require a new persistent service.
A practical first step is to choose five to ten current files and 30–50 questions from one process. Test ordinary heading and keyword search first. Add passage retrieval and generation only where they save time without creating unsafe answers. Display a link to the approved procedure version and let staff open the original. The value of this research case lies in its restrained role for the model: it can help find and explain a rule, while people remain responsible for the rule and the action.
