Not every email should start with a search of the entire knowledge base
For a small company developing ERP and procurement systems, support can quickly become a sorting center. Simple questions, faults, and issues requiring an engineer arrive in the same inbox. If staff manually determine urgency, find the right instructions, and write every reply from scratch, complex cases wait longer than they should.
Croatian company RIS d.o.o., listed as having 10–49 employees by the European Digital Innovation Hubs Network, tested a different process with EDIH Adria. Its pilot assistant classified incoming emails, searched FAQs, manuals, and examples of previous requests, and suggested response drafts. A support agent could edit and approve them. That is different from a public bot answering without review.
The project was reported in 2025; we examine it as a practical case, not today's news. The boundary between a demonstration and production matters most. The EDIH Network reports estimated reductions in workload and response time, but the project's organizer explicitly calls the solution a functional prototype tested on simulated real-world scenarios in a limited environment. Public sources do not disclose sample size, measurement methodology, or sustained operation on live traffic.
What was actually built
According to EDIH Adria, RIS worked with a small internal knowledge base of FAQs, software instructions, and example inquiries. The assistant demonstrated four functions: email classification and prioritization, retrieval of relevant documents, proposed answers under agent supervision, and an interface showing interaction history. The fuller EDIH Network case names RAG and LLMs, Python and Streamlit for the demonstrator, and IMAP integration with the company mailbox.
The architecture separates different jobs. Classification helps decide who should see a request and when. Retrieval returns evidence from approved documents. A generative model composes a convenient draft. Finally, a person checks facts, tone, and the right to disclose information. If these steps are compressed into “AI answers by itself,” locating an error becomes difficult.
The sources do not identify the exact model, its hosting location, or infrastructure costs. Do not describe this pilot as an on-premises deployment. A Russian small business could separately consider a local option when email and knowledge data must stay inside a controlled environment; that is an editorial recommendation, not a reported RIS property.
Reading the reported outcomes
The EDIH Network cites an estimated 50–60% reduction in manual effort on simple requests, more than 70% faster replies to typical requests, and up to 40% fewer escalations to senior support. EDIH Adria's own page clarifies that these are potential impacts assessed through tests and simulations. The figures cannot be inserted into a budget as verified production KPIs or turned into a promise of round-the-clock service without further validation.
The baseline reply time, number of emails, classification accuracy, and work spent correcting drafts are not disclosed. Neither is a financial return on investment. The workflow is still informative: the pilot kept the agent responsible for sending responses and produced technical documentation and a roadmap before production integration. Its result was a test of the design and requirements, not demonstrated replacement of the support team.
The project report includes a crucial qualification: the knowledge base must be current and structured, and integration with CRM and ticketing should be planned early. Otherwise, assistance in one window creates manual copying into another. The robot may enthusiastically sort every email into colorful piles, but it should not own a customer's case without a process owner.
Adapting the principle for a Russian company
Start with one stream, such as the support inbox for one product. Take 100–200 anonymized past emails and label them manually: request type, urgency, correct specialist, permitted source document, and final answer. Remove personal information and confidential attachments from the test set or keep them entirely within a protected environment. For questions without an approved answer, record that escalation to an engineer is the correct outcome.
A pilot workflow could be:
- A mail connector reads copies of incoming messages without send or delete rights. Reprocessing one email must not create another ticket.
- Rules and a classifier suggest a category and priority; high-risk matters—payments, access, outages—go directly to a person.
- RAG searches only approved versions of manuals accessible to the relevant agent and customer. Drafts display citations to specific sections.
- The operator edits and sends through the existing support system. The system then records the time and nature of edits for evaluation.
This is not a description of RIS's exact implementation. Some elements, including read-only access and duplicate control, are proposed safeguards for a new deployment. During a pilot, do not give the model permission to send emails on its own. An incoming email is untrusted text: an attacker may include instructions to ignore rules and disclose another customer's data. OWASP documents such scenarios for assistants and notes that RAG alone does not eliminate prompt injection.
Data, environment, and limits
The key asset is not the LLM but a maintained knowledge base. Every manual needs a product, version, effective date, and owner. Historic correspondence can contain outdated promises or customer-specific contract terms; do not convert it blindly into reference truth. Some routine questions are better answered by templates and rules without a generative model. Where the answer depends on an order or contract status, the system needs an authorized query to the system of record, not a model's guess.
An on-premises deployment would need a mail gateway, document and log storage, a search index, a draft-generation model, and access controls. Hardware sizing cannot be inferred from the RIS case because its model and load are undisclosed. Measure your own stream: emails per hour, attachment lengths, search time, and acceptable latency. A cloud pilot may cost less if data transfer is permitted and volume is low; otherwise, evaluate a local stack including its full ownership cost.
For personal data, set retention periods, mask test data, and apply role-based access. Check that drafts do not copy details from another customer's ticket or change the meaning of an instruction. Leave complaints, safety issues, money, and contractual commitments to specialists. “Draft ready” is not “approved for sending.”
Economics without substituting someone else's estimate
An illustrative calculation, not an RIS outcome: an inbox receives 400 emails per month, 60% of which are repetitive. If an assistant saves four minutes on half of those repetitive emails, that is eight hours per month: 400 × 0.6 × 0.5 × 4 / 60. Freed time is not profit. Review time, knowledge-base work, API fees or local hardware, maintenance, and error costs must still be subtracted.
Compare a control group and the pilot on the same request categories. Core measures include time to a useful first answer, operator work time, correct routing, drafts with verifiable citations, material edits, and escalations. Count false priority separately: if AI marks every email urgent, the queue has not become smarter. Count a result only after an answer has been sent and accepted, not merely after a draft is created.
A practical next step
Run a two-week shadow test on one inbox. In week one, gather and label inquiries, clean the knowledge base, and approve escalation rules. In week two, show assistant suggestions to agents but prohibit automatic sending. Record the correct category, retrieved document, edits, and time for each email. If there are too few repeatable requests, do not expand the project for the sake of a polished demo.
The RIS case is valuable not as a promise of “60% less work” but for its authority boundary: classification and retrieval accelerate preparation, a person controls the reply, and impact is first measured on the company's own traffic. That makes it easier for a small business to decide whether it needs a local model, RAG, or simply more discipline in its knowledge base.
