Why a price per million tokens is misleading
With an API, the bill generally rises with actual consumption. With a self-hosted server, a substantial share of costs remains even when no one sends a request: hardware or rental, administration, redundancy, monitoring, and upgrades. Dividing a GPU's monthly cost by its maximum advertised throughput is therefore usually a poor estimate for a small business.
In research published in June 2026, Chitral Patil estimated inference costs at different request rates, with different models and hardware. On the same H100, the author found a wide range of effective prices as utilization changed. This is neither a universal price list nor a forecast for a Russian company: it illustrates why assuming continuous full utilization can significantly understate self-hosting costs. The author also repeated the core sweep on an A100 and noted that results depend on the hardware platform.
For a manager, the better question is: how much does a **successfully completed business task** cost under our actual arrival pattern? Tokens are useful for technical accounting, but they do not price an idle server, manual corrections, or a late response to a customer.
Measure the workload profile first
Collect at least two to four weeks of real or de-identified requests. For each, record arrival time, input and output tokens, response duration, task type, quality result, and whether a person had to intervene. For RAG, account separately for retrieval, document preparation, and text extraction; they are not free add-ons to generation.
A daily average is not enough. Ten thousand requests spread evenly across a month and the same ten thousand concentrated into short business-hour peaks require different capacity. Examine the hourly distribution and at least the p95 response time during peaks. A support workflow may need a strict time-to-first-response target; overnight document processing may tolerate more delay.
vLLM exposes metrics for queued and running requests, time to first token, end-to-end latency, and processed tokens. Its `bench serve` utility can replay a load profile; a test with one short prompt does not represent real context lengths or simultaneous-request spikes. NVIDIA Triton's documentation likewise describes the throughput–latency trade-off introduced by batching. These are engineering instruments, not a guarantee of a particular saving on your hardware.
Illustrative calculation: the break-even point
Assume a company compares two approaches for the same validated task. All numbers below are **hypothetical**, in roubles; they are not supplier prices or a market estimate:
- self-hosting: RUB 90,000 of fixed monthly cost, including depreciation or rental, operations, and reserve capacity, plus RUB 0.15 of variable cost per processed request;
- external API: RUB 1.20 per request at the actual mix of input and output tokens;
- quality, availability, and human-correction rates are initially assumed equal; tax, currency, and contract details are excluded.
Self-hosting then costs `90,000 + 0.15 × N`, while the API costs `1.20 × N`, where `N` is monthly request volume. The lines cross at about **85,700 requests per month**. At 20,000 requests the model yields RUB 93,000 versus RUB 24,000. At 100,000 it yields RUB 105,000 versus RUB 120,000. That is arithmetic based on assumptions, not a savings promise: the local system must still handle the actual peak load within the required latency.
Now assume the local model needs a manual correction two percentage points more often, and each correction costs RUB 20. That adds RUB 0.40 per request. All else equal, the threshold moves to roughly **138,500 monthly requests**: `90,000 / (1.20 − 0.15 − 0.40)`. A small quality gap can outweigh cheaper inference. If the denominator becomes zero or negative, volume growth alone never makes the local option cheaper in this model.
Compare more than roubles per token. Measure the cost of an accepted result: how many requests were handled correctly, how many documents needed review, and how many employee hours were genuinely saved. Different models or review policies can move total cost in either direction.
When a local model makes sense without direct savings
Restrictions on data transfer, isolation requirements, dependence on an external connection, and cost predictability may independently justify a local deployment. But “local” does not automatically mean secure. It still requires access control, upgrades, access audits, backups, log governance, and rules for documents ingested into RAG.
A small company can also consider a split arrangement: retain sensitive steps in a protected environment while routing permitted, infrequent tasks to an API; or use existing compatible infrastructure instead of buying a dedicated always-on GPU. Such a split requires checking contracts, data categories, request routing, and consistent quality criteria for both paths. If data cannot be sent outside, an API is not an admissible option whatever its price.
A low-commitment test
1. Choose one process with a measurable outcome, such as a knowledge-base answer draft or a standard document review. Agree on a reference set and the cost of mistakes.
2. Collect request timing and real context lengths. Remove personal data from the test set or run the evaluation within an approved environment.
3. On comparable data, measure the API and local candidate: accepted-answer rate, manual corrections, peak-hour p95 latency, queue size, and actual throughput.
4. Include the full fixed monthly cost: hardware or rental, power, administration, redundancy, monitoring, and refresh cycle. Count variable spending and human-review costs separately.
5. Decide using observed volume and two growth scenarios, not the maximum speed from a demo. Update the model with actual metrics after the pilot.
The practical conclusion: a local model becomes economically interesting not when it reaches an impressive tokens-per-second figure, but when utilization, quality, and operating cost together deliver a cheaper **accepted** result at the required latency. With sparse, spiky traffic, a simple per-token comparison often conceals the most expensive line item: the hours the server waits.
