Agree on a deadline before buying speed

A small company often plans a local model as a 24-hour chat service: every request must be answered immediately. But some work has a different shape. Draft replies to yesterday's emails, request classification, catalog checks, and duplicate detection may only need to be ready by morning. For these tasks, the relevant requirement is the batch completion deadline, not the first second of response. That changes both architecture and capacity planning.

Consider a team handling several thousand requests per month. Some messages are urgent: a customer is waiting for an operator, an order is changing, or an escalation is needed. Others are overnight draft preparation for a specialist. If everything enters one queue without rules, a long batch may delay an urgent request. If a powerful node is kept running around the clock solely for rare peaks, capacity may sit idle. The aim is not to “move everything to the night” but to assign each type of work an appropriate deadline.

Two different meanings of batching

A scheduled batch is a business timetable: collect jobs by a cutoff, process them, and deliver results for review before the next shift. The queue retains a job ID, input-data version, priority, and deadline. This mode is possible even without a special request-combining mechanism inside the model server.

Dynamic or continuous batching is an inference-server mechanism. It combines compatible requests during computation to improve throughput. NVIDIA Triton documentation notes that waiting longer to form a batch may improve throughput at the expense of latency; queue settings and timeouts are configurable. For generative models, llama.cpp documents continuous batching and parallel server slots. The Orca research paper explains why scheduling at generation-step granularity can be preferable to inflexible waits for an entire batch to finish. These are engineering capabilities, not promises of savings on your hardware.

Scheduled batches and server-side batching can be used together, but they solve different problems. The first relaxes the business requirement for an immediate response. The second may make better use of hardware when several requests are active at once. With very little traffic, there is simply nothing for the server to combine. If latency is critical, waiting for a batch to fill is not acceptable.

A small-business architecture

At intake, split work by explicit rules rather than a model's guess: urgent messages go to an interactive lane; mass draft preparation goes to a deferred queue. Each job needs a unique ID, source, owner, deadline, template version, and permitted data scope. Remove unnecessary personal information before queuing; an on-premises server does not waive internal access rules.

A scheduler takes jobs in groups and calls the model. Output is saved as a draft, not sent to the customer automatically. Deterministic checks reject missing fields, impossible categories, stale source records, and duplicates. A specialist approves or corrects the result. Only then does the integration layer record the approved action in the CRM or ticketing system. The audit trail joins the source request, model and prompt versions, response, human decision, and final status.

The queue itself needs safeguards: maximum size, waiting-time limit, retry count, and a separate error list. A huge document or stuck request should not occupy all capacity and block other work. When a deadline is exceeded, hand the job to a person or a fallback process; a silently growing backlog turns savings into missed commitments. For sensitive actions such as refunds, payment-detail changes, or promises to customers, the decision remains with an authorized employee.

Count the economics honestly

Measure the cost of an accepted result delivered on time, not just tokens per second. Include infrastructure, queue storage, node startup and shutdown, integration, quality control, human review, errors, and headroom for a peak day. Track the share of jobs completed before their deadline and the latency of urgent requests separately.

An illustrative model with explicit assumptions: 6,000 similar requests in a 30-day month, or 200 per day on average. A pilot on your own inputs found that a particular model handles a day's batch in two hours including checks and headroom. That is 60 hours of active work per month. An always-on node is available for 720 hours in the same 30 days. If compute is billed hourly and the resource can actually be shut down, compare 60 times the hourly rate with 720 times the rate, adding startup time, storage, and support. The 660-hour difference illustrates only the infrastructure-hour component; it is not a forecast of total savings.

If a physical server has been purchased and stays on for other work, moving jobs overnight does not recover capital expenditure by itself. It may free daytime capacity and reduce contention with interactive requests, but the result depends on utilization, electricity, and alternative uses of the node. If the actual batch takes six hours but must finish in three, the two-hour assumption is worthless: change the model, parallelism, or deadline. Larger batches may also worsen latency, memory pressure, or reliability. Choose settings by measurement, not a universal number.

What to measure in a week

Choose one process with a movable deadline and a real input sample: short and long requests, attachments, duplicates, and malformed items. Compare three modes at equal input and quality: one-by-one processing on arrival, small groups during the day, and one deferred batch. Measure time to first response for urgent jobs, on-time completion for deferred jobs, jobs per hour, peak memory, accepted-draft share, and employee review time.

Do not batch tasks merely because “night is cheaper”: customer promises and operational consequences matter more than GPU utilization. A sound first step is to define two queues and a completion deadline for each, then run a shadow pilot without automatically sending outputs. Expand if the batch meets its deadline and lowers the cost of an accepted task. Otherwise, ordinary routing or a template may solve the problem without another server. The robot intern will gladly carry the whole tray of letters at once; a manager should still keep the red urgent folders on the desk.