Where a business pays for the same work twice

Consider an internal assistant that answers questions about a sales policy. Each request begins with the same system instructions and a long policy excerpt; only then does it add the employee's specific question. A local model must process all that input before it can produce the first word. When the common part is long and unchanged, this repeats computation. It affects latency and the number of requests the existing server can handle.

vLLM provides Automatic Prefix Caching (APC) for this pattern. The engine keeps the KV-cache for previously processed blocks and reuses it when a new request has the same beginning. This is neither a stored answer nor document search. The model still processes the new question and generates a fresh response. The mechanism saves work on the common initial portion of the prompt—the prefill phase. vLLM's documentation explicitly notes that prefix caching does not speed up the generation of new tokens, known as decode.

For a small company running AI on premises, this may be more valuable than buying a larger GPU if one model handles many short requests based on stable templates: classifying enquiries, filling CRM fields, explaining returns rules or answering repeated questions about a manual. But the first step is to measure how much text the real workload shares, not merely to enable a setting and announce savings.

What has to match

The important word is “prefix”. Shared content must appear at the start, in the same order, and produce matching blocks after tokenisation. If the application inserts today's date, a random identifier, a user name or a changing tool list before the shared instruction, reuse may stop within the first blocks. If a RAG pipeline selects and orders documents differently on each request, caching a long common beginning also becomes harder.

A practical pilot architecture is to:

  • keep stable system instructions and unchanged rules at the beginning;
  • put variable data—questions, customer records and recent entries—after the common block;
  • version the template so a rule change deliberately creates a new prefix;
  • keep RAG where answers depend on an evolving knowledge base: caching is not a replacement for retrieval;
  • confirm that application message ordering and serialisation remain stable across requests.

Do not impose this structure at the expense of response quality or access rules. The common block is only the portion that can legitimately be shared across requests. Personal or confidential information must not be placed into a “shared instruction” just to improve the cache-hit metric.

A model calculation, not a payback promise

Suppose a working day contains 1,000 requests. Each has 2,000 tokens of identical instructions and reference text, followed by 300 tokens of unique question and context; each answer averages 200 tokens. These are illustrative assumptions, not measurements from a company. Without prefix reuse, the engine processes 2.3 million input tokens: 1,000 × (2,000 + 300). If the shared prefix stayed in cache and every later request hit it, repeated input processing would approach 302,000 tokens: one processing pass over 2,000 shared tokens plus 1,000 × 300 unique tokens. The difference is about two million tokens of prefill computation.

That does not mean a bill or response time falls by 87%. The token arithmetic indicates potentially avoided work, not a monetary tariff. Generation of 200,000 output tokens remains; the cache occupies memory; the first request warms it; some requests will miss; and the server may spend much of the day waiting for people. With a fixed-price GPU lease, the gain becomes financial only if the organisation can actually reduce rental hours, use a smaller configuration, postpone expansion or complete more useful tasks on the same hardware. For an owned server, electricity, maintenance and reserve capacity matter too.

The illustrative ratio therefore should not be copied into a budget. PrefixBench-H100, an independent research benchmark released in September 2026, varies shared-prefix length, suffix diversity, concurrency, answer length and cache pressure. Its authors identify conditions in which time to first token improves substantially and conditions in which memory pressure erodes the benefit. This is a measurement on an H100 under specified conditions, not a performance guarantee for your server.

How to measure the effect on your workload

Take a de-identified one-week sample: at least several hundred real requests from one workflow, including busy and quiet periods. If data must not leave the local environment, collect statistics inside it and do not export the requests themselves for analysis. Separate stable-policy questions, long conversations and diverse one-off tasks. Measure the length of the repeated initial block after applying the same tokeniser as the production model.

Then run two tests on the same model and hardware, with and without APC. Keep request arrival order and rate, concurrency limits and output lengths comparable. A test that fires every request at once does not resemble a working day: a warm cache in a compact benchmark may look better than a traffic pattern with long gaps and block eviction. Account for warm-up, restarts, prompt revisions and whether returning requests reach the same server replica.

vLLM exposes counters for prefix-cache queries and hits, KV-cache block usage, time to first token, prefill time and end-to-end latency. Read them alongside the number of useful completed tasks per hour and the proportion of answers a human accepts without correction. A faster bad answer is not a saving. Use the hourly cost of the actual configuration, rather than an abstract token price, to compare the cost per accepted task in both modes.

A useful decision rule is this: if time to first token falls, peak throughput increases, quality is unchanged and cache memory does not crowd out other required requests, the cache has practical value. If answers are mostly long, prompts have no common start or the server is rarely busy, optimising APC before the data and the workflow is premature.

Limits and security

The KV-cache competes with active requests for memory. Under pressure, blocks may be evicted and expected hits disappear. In a distributed setup, routing also matters: a repeated request sent to another server instance without the same state cannot use the first instance's local cache. No cache-hit percentage, by itself, guarantees lower latency when queuing or contention changes.

There is a security boundary around a shared cache. vLLM's security documentation describes a timing side channel through which someone on a shared backend may infer whether another user's prefix was present. For a multi-tenant setup, vLLM provides `cache_salt`: different secret salts isolate cache sharing between trusted groups. This also reduces potential reuse and therefore some of the performance gain. The salt should not be a predictable customer identifier; the security owner must define the isolation scheme. A single-tenant local deployment may not need salting, but access control for the server itself remains essential.

Start with one workflow and one template. Over a week, measure shared-prefix frequency, cost per accepted answer and the difference between the two modes. Only then decide whether a larger server is needed and whether prompt construction should change. Caching prevents repeated computation; it is not a universal discount on local AI.

Server-room photo: Carl Lender, CC BY 2.0; centrally cropped to a square by the editors. The original file page and licence are linked in the sources.