What the case actually shows
Mentana is an Australian startup selling a platform for supply-chain operations. According to the Google Cloud customer story, it combines information from ERP, CRM, and point-of-sale systems. Gemini-based agents parse incoming emails, check inventory in ERP, and suggest purchasing or production adjustments. Another workflow investigates discrepancies between physical stock and digital records. Mentana's own website describes stocktake investigation and stock-transfer requests that a user can initiate with a click.
This is a vendor serving large enterprise clients, not a verified outcome from a Russian small business. The transferable lesson is therefore the sequence of work, not its scale: bring together evidence from several systems, identify a specific exception, explain it, and let a person decide whether to change the record. The robot may be diligent, but its digital smile is not a signature on a stock adjustment.
The Google Cloud story reports a 22% reduction in infrastructure costs after a cloud migration, 99.99% system uptime, and a move to fortnightly release cycles. These are Mentana infrastructure metrics, not agent accuracy, shrinkage reduction, or agent ROI. The public sources do not provide an error sample for the agents or the cost per successfully resolved stock exception.
A suitable task for a smaller company
For a retailer with one or a few warehouses, a sensible starting point is investigating a discrepancy such as “the system says it is available, but the shelf is empty,” or the reverse. An operations manager normally compares exports, looks for order reservations, receipts, returns, transfers, and unit-of-measure mistakes. If relevant evidence is spread across supplier email, the accounting system, and a stocktake spreadsheet, finding the cause can take longer than correcting it.
AI is not a replacement for the system of record. Its useful role is to read unstructured correspondence or a stocktake report, match product references, form a hypothesis, and point to the underlying records. Exact arithmetic, duplicate detection, and permitted-action rules remain conventional software. RAG is useful only when the explanation requires retrieval of a receiving policy, returns procedure, or supplier terms. There is no reason to build document retrieval merely to compare two numbers in a table.
For example, an assistant might say: “The count sheet records 12 boxes; the ERP has 120 units, and the product card specifies 10 units per box. This could be a unit-of-measure issue; verify the delivery note and batch.” This is an editorial example, not a quote from Mentana's deployment and not a claim about accuracy. If the ERP already handles unit conversion reliably, a deterministic rule is the better tool.
A minimal implementation path
A pilot does not require giving an agent unrestricted permission to edit stock. Start with one product category and one exception type. A practical flow could be:
- A service ingests stocktake reports and related documents from approved sources. Incoming email is treated as data, never as an instruction to the agent.
- An integration reads the SKU record, unit of measure, inventory at a specific time, reservations, recent receipts, and movements from the ERP. Each record retains its identifier, timestamp, and source.
- Conventional software reconciles the figures and filters cases explained by known rules: reservations, synchronization delays, or packaging conversions.
- The model examines only the remaining exceptions, proposes a possible cause, and assembles a short explanation tied to source records.
- An accountable employee accepts, rejects, or refines the conclusion. Stock adjustments, purchases, transfers, and messages to suppliers require a separate approval step.
- The system records the decision, reason, and action trail. Reprocessing the same report must not create a second adjustment.
Each system needs a defined data owner. If the point-of-sale system records a sale, the ERP records a reservation, and the stocktake sheet records a physical count, decide in advance which field answers which question. The newest record is not always the truth: sales and returns may reach the ledger late. Ambiguous cases need an “insufficient evidence” state, not a confident guess.
Cleaning the product master can matter more than model selection. The same item may appear under a supplier code, internal SKU, and abbreviated name; boxes and individual units may be mixed. In the pilot, create a separate mapping of aliases, units, and batch-matching rules. Without it, the assistant may persuasively discuss the wrong product.
Local models, cloud services, and access boundaries
The documented Mentana setup uses Gemini and Google Cloud infrastructure. Mention of hybrid or on-premises platform deployments does not establish that Gemini itself runs locally. For a Russian company, a local model is a separate architecture option, not a fact about Mentana. It can make sense when stock data and internal documents may not be sent to an external provider, workload justifies owned hardware, or a predictable security boundary is required. If the task is infrequent and the contract permits external processing, an API may be simpler and cheaper. Test the choice on actual documents and include full operating costs.
In either case, limit the assistant's tools to reading specific tables and documents. An analysis service account should not be able to change inventory. A separate service can perform writes only after explicit human approval, validated fields, and an idempotent operation identifier. OWASP calls unnecessarily broad agent permissions “excessive agency” and recommends least privilege and human approval for consequential actions. Supplier email is untrusted external input and may contain an attempt to redirect the agent; never treat it as an authorized command.
Logs should retain enough for an audit—source-record references, rule version, model conclusion, approver, and action—without copying every commercial record. Define retention periods and role-based access. If the model or integration fails, route cases back to the manual queue instead of allowing the system to invent inventory numbers.
Economics without borrowing someone else's percentage
Measure the cost of one correctly resolved exception in your own company. For two weeks, collect a baseline: case volume, investigation time, repeated cases, the cost of delayed orders, and the cost of a mistaken adjustment. Compare the same stream during a pilot, including review time, integration work, model operations, and error correction. “The agent read a hundred emails” says little if half of its conclusions had to be redone.
Here is an illustrative calculation, not a Mentana result or a reader forecast. Assume 200 exceptions per month, 15 minutes of manual investigation each, and a fully loaded staff cost of 800 rubles per hour. Current labor is 40,000 rubles per month. If a suggestion saves an average of eight minutes in only 70% of cases, it releases about 18.7 hours, or approximately 14,900 rubles per month before implementation and support costs. If the pilot and maintenance cost more than the saved effort and avoided mistakes, there is no reason to deploy automation at this scale.
Measure quality separately: the share of suggestions accepted unchanged, false SKU matches, ERP changes reversed after review, and time to close a case. With a small volume and simple rules, conventional reconciliation may outperform an agent in both cost and reliability. That is also a valuable pilot result.
The next step
Take 30–50 recently closed discrepancies of one type. Reconstruct the original documents, product movements, and final human decision for each. If a meaningful share cannot be explained without interpreting correspondence and business rules, test a read-only assistant. Set a quality threshold and a ceiling on cost per accepted resolution in advance. Add write permissions only after demonstrating repeatable quality, traceable explanations, and a working approval process.
The practical lesson is not that every warehouse needs a cloud agent. It is that an agent can help where fragmented evidence meets operational exceptions, while responsibility for physical stock and ledger changes remains with the business.
