Why an ordinary chatbot was not enough

i-Tom Solutions operates FGO, an invoicing platform with related accounting workflows. According to a case study published by AWS, it serves more than 170,000 small and medium-sized customers and received 200 to 600 written inquiries per day. Some questions recurred, while others called for technical investigation or an examination of a customer's specific circumstances. If specialists spend their time repeating standard explanations, complex cases wait in the same queue.

The company's first goal was narrow: answer routine questions outside business hours. This is an important boundary. An assistant should not pose as an accountant, alter an invoice, or issue a personalised judgement based on one message. It should find current instructions, draft a response, and pass complex cases to a person.

AWS presents this as its own customer case study, not an independent audit. The figures below are therefore statements from the company and infrastructure vendor. A public inquiry dataset and accuracy evaluation protocol are not provided.

Otherwise, a bot can work a very polite night shift and leave the morning team to clean up the consequences.

What the team built and how it launched

i-Tom worked with AWS solutions architects and partner Auvaria on a four-to-five-week proof of concept. The design used RAG: search a centralised knowledge base, then generate a response using the retrieved context. The published architecture names Amazon Bedrock for generation, AWS Lambda for document processing, and Amazon OpenSearch Service for semantic search. After the proof of concept, the team deployed the validated configuration to production in one week using infrastructure templates.

The sequence matters more than the cloud vendor:

  • collect current instructions and frequent questions, assigning an owner to each section;
  • establish ingestion, updates, and deletion for material in the search index;
  • define which question types may receive an automatic answer;
  • show operators the retrieved sources and the reason for escalation;
  • expand access to customers gradually while reviewing responses and the error log.

According to AWS, the assistant is currently in a limited production pilot with selected customers. The company says it handles hundreds of inquiries daily. The account does not establish that all of them are resolved without a person, that 600 inquiries is a constant daily load, or that the assistant covers all support work.

Transferring the approach to a Russian business

A deployment within a Russian company's own environment does not have to replicate AWS services. A local document parser, search index, and self-hosted model can fill the same roles. But “local” does not automatically mean “safe”: document permissions, the routes taken by personal and commercial data, conversation retention, and operator access still need review. Model and hardware selection should follow measurements of real traffic and Russian-language answer quality.

Start with a low-risk channel, such as questions about navigating the product or using a standard feature. Leave disputed tax interpretations, individual amounts, invoice operations, and changes to customer records outside the first wave. In those cases, the assistant can prepare a draft or collect relevant information, but an employee should decide. When no source exists or documents conflict, the correct outcome is to avoid a confident answer and hand the case to a person.

A minimal technical flow has an incoming queue, intent classification, permission-aware retrieval, a source-linked answer draft, escalation rules, a decision log, and an operator screen. The CRM or ticketing system remains the system of record. The model should not have direct permission to modify financial records. Round-the-clock service also requires monitoring, load limits, and a fallback route when the model or index is unavailable.

Economics: count safely resolved questions, not messages

The public case study does not disclose infrastructure spending, maintenance costs, the wrong-answer rate, or a measured reduction in staff effort. Its figures cannot support a payback promise. For your own pilot, calculate the cost per accepted resolution:

**Total monthly costs / number of questions the assistant resolves correctly without rework or a repeat contact.**

The numerator should include more than compute: knowledge-base preparation, quality review, integration maintenance, operator corrections, and incidents. Track escalations and repeat contacts separately. A chat that rapidly produces inaccurate answers may improve first-response time while increasing the burden on the second line.

Here is an illustrative calculation for pilot selection. Suppose 6,000 written questions arrive in a month, 40% are routine, and 70% of routine answers pass review. That would leave 1,680 potentially resolved without rework. These are not i-Tom results or a forecast for another company; each share needs measurement on your own inquiries. At total monthly spending of RUB 100,000, the illustrative cost would be about RUB 60 per accepted answer. Compare it with human handling of a comparable question and the cost of a mistake.

Checks before expansion

Collect 100–200 anonymised real questions, including ambiguous cases and those that must not be answered automatically. For each, record the correct source, acceptable response, and handoff condition. Run the assistant in shadow mode: it suggests an answer, but the customer only receives a human-approved message. Measure correctness, the completeness of document references, confidently wrong answers, time to resolution, and repeat inquiries. AWS's RAG evaluation documentation includes citation coverage and precision checks; these measures are useful regardless of platform.

The next step for a manager is to select one routine question category, name an owner for the knowledge base, and agree on a threshold at which automatic answers are switched off. If the test cannot demonstrate value after corrections are counted, it is too early to widen the channel. The main lesson of i-Tom's story is the controlled sequence: a bounded proof of concept followed by a limited launch, not the purchase of a supposedly universal AI solution.