What changed
vLLM released version 0.31.0 on 5 October 2026. Its release notes identify two changes relevant to operators of locally hosted AI. The new `--max-num-active-seqs` setting caps requests in the RUNNING state independently of the existing `max_num_seqs`. Also, per-request `mm_processor_kwargs` and `media_io_kwargs` are now rejected unless the server is started with `--trust-request-mm-kwargs`. That behavioural change may affect clients that configured image or audio processing in individual requests.
Why a business should care
When support, sales and internal search share one inference server, a burst of long requests can increase waiting time for everyone else. The new cap lets an operator control admission into execution separately. It is not, however, a ready-made priority policy for departments or an overall inbound traffic limit: queueing, timeouts and rejection rules still need to be configured and measured. Google's independent SRE guidance notes that requests of different sizes consume different amounts of resources, so requests per second alone is a poor measure of service capacity.
The tighter default for multimodal options can reduce exposure to untrusted request-level controls, but upgrading may break an existing integration. Test actual application traffic, not merely whether the model starts: images, audio, long context and concurrent users all matter.
An important limitation
The release also lists optimisations for particular large models and accelerators. They cannot be generalised to every local server or treated as a promise of lower costs. First compare time to first token, queue wait, error rates and memory use on your own model before and after upgrading. The official vLLM illustration shows a four-GPU architecture as an example; four GPUs are not a requirement of version 0.31.0. Primary source: the vLLM release notes.
