The local perimeter starts before the first request

When a company says it runs AI locally, it usually means customer documents are not sent to an external API while the model works. That is important, but it leaves another question unanswered: what entered the server during installation? Weights, configuration, tokeniser, adapters, model code and dependencies all arrive from outside. Downloading them is a software supply-chain decision. OWASP identifies third-party pretrained models, data, adapters and serving platforms as parts of that supply chain.

A small business does not need a dedicated model-security research team to address this. It needs a clear intake route for one artifact: who published it, what use is permitted, which exact version was tested, what files are loaded and who approved its move into production. The route matters especially when the model will see contracts, customer requests or internal knowledge bases.

Three different questions that are often confused

The first is whether a file can execute code while it is loaded. Hugging Face explicitly warns about pickle deserialisation; many older PyTorch checkpoints use this format. `safetensors` stores tensors without pickle's arbitrary instruction-execution mechanism and is preferable for weights. But a safe format for one file does not make the rest of the package safe. If starting the model requires `trust_remote_code=True`, the application is trusting custom code in the model repository. Review and approve that code separately, or choose a model supported by the installed library without remote code.

The second is integrity and provenance. A SHA-256 digest shows that a local file matches a chosen reference; it does not prove that the publisher is trustworthy. If both the digest and the file come from the same compromised page, a match provides little assurance. Record the developer's official organisation, the original repository address, the full commit identifier and your own file-hash manifest after intake. Hugging Face Hub supports downloading a snapshot from a specific revision; its documentation says a hash revision requires the full commit hash, not a shortened one.

The third is fitness for purpose. A model may be honestly published and use a safe file format, yet still be unsuitable for a commercial workflow because its licence is unclear, its use is restricted, its Russian-language performance is inadequate, its memory requirements are too high or it fails on your documents. A file format is not a quality certificate, and an antivirus badge does not establish commercial usage rights.

An intake route without heavy bureaucracy

For one pilot, a supply record in an ordinary internal register is enough. It should contain:

  • the original publisher and exact repository URL, rather than a similarly named third-party quantisation;
  • the model card, licence and commercial restrictions, including any adapter's terms;
  • the full commit of the chosen revision, the list of files actually needed and their checksums;
  • the serving engine, library, tokeniser, chat template and dependency versions;
  • memory requirements, supported environment and allowed network access;
  • the review owner, approval date, test results and rollback procedure.

Download the package first into a separate area without production documents or secrets. Do not run arbitrary installation commands copied from a README on a production server. Filter the file list where the format and runtime permit it: the Hub client library provides `allow_patterns` and `ignore_patterns` for this purpose. Filtering does not replace inspection of configuration and dependencies. If a loader asks to execute remote code, treat that as a separate approval decision rather than a convenience checkbox.

Look at the platform's scan results, but understand their limits. Hugging Face says it triggers malware scanning for repository files on each commit; a missing badge may indicate that scanning is pending or has failed. Its pickle documentation also makes clear that import checks and related lists are not a complete package audit. “The hub did not complain” is a useful signal, not permission to run a download with access to a production database.

After the initial review, start the package in a restricted environment: without outbound connections unless they are required and approved; without CRM credentials or real personal data; and with a small set of synthetic test queries. Check which files the process actually reads, whether it fetches additional dependencies at startup, how much memory it needs and whether it passes controlled examples. Then copy exactly the approved file set to an internal artifact store and let the production service load models only from there.

This internal gate need not be an expensive new platform. For a small deployment, it may be a protected directory or existing storage where only one responsible role can write and changes are logged. What matters is separation of duties: a developer proposes an update, a reviewer approves the package, and the production service reads only an approved version. Do not rely on a floating `main` reference in production configuration; the same address may refer to a different set of files tomorrow.

Check behaviour separately from files

Neither a checksum nor `safetensors` tells you how a model behaves in the business process. OWASP also discusses modified or undesirable weights and weak provenance assurances. Add a short task-specific evaluation to technical intake: 20–50 typical and edge-case queries from the chosen workflow, de-identified or synthetic. Include Russian phrasing, abstention when information is missing, source references for RAG and cases that must go to a human decision-maker. Compare the results with the previously approved version if one exists.

This is not “certification of model safety” and will not reveal every hidden property in the weights. It is a minimum check that replacing the artifact did not degrade the business function you need. Track prompt and knowledge-base versions separately: if they change at the same time as the model, the cause of any regression is hard to identify.

The cost of control and a concrete first step

The main cost of this route is a responsible person's time and space for one verifiable copy, not a new GPU. For a single model updated occasionally, a supply record and short test are usually less costly than finding out after deployment why the weights changed or who gave downloaded code access to business data. This is a management comparison, not a calculated return on investment for a particular company: actual effort depends on the number of models, update frequency and isolation requirements.

Take one model your team intends to deploy locally this week. Ask for the official source, licence, full commit, file list, whether `trust_remote_code` is needed and the rollback path. If any answer is unknown, do not move the package into production until it is resolved. Local execution protects data during inference only if the perimeter and the artifacts entering it are managed.