What changed

The Xinference team released version 3.5.0 on September 25, 2026. Its release notes list model request logs, correlation identifiers that connect events, a separate log view in the interface, and protected retrieval of request bodies. The version also adds recommendations for model launch settings and model sizes based on local memory metadata. This is an inference-server update, not a new release of base-model weights.

For a small IT team, the practical value is not another chart. If a RAG assistant returns a wrong answer, operators can connect the user request to the model call and system events, locate the failure, and repeat the check. The implementation notes describe `request_id` and `correlation_id`; a separate `model_requests:read_body` permission restricts access to request contents.

The important limitation

Detailed logging is opt-in and disabled by default. The developers warn that request values can contain sensitive data. Their implementation notes describe a limit on stored body size, file rotation, and `0600` permissions, but these controls do not replace a company's decisions on retention, administrator access, and excluding secrets from requests. NIST's generative AI profile also identifies data privacy as a material risk.

A Russian business already serving a local model for an internal knowledge base or an agent could first enable correlation without retaining bodies, check access controls, and only then decide whether full logging is needed for a short test set. The release provides no public measurement of staff time saved or proof that answer accuracy improves: these are observability capabilities, not a demonstrated business outcome.

The image is an official screenshot of the Xinference model-launch interface from the project's repository. It shows the product, not the new request-log page. The screenshot is distributed with the repository under Apache License 2.0; only a square crop was made, without generative alteration.