← The journal
Security Guide

Model Supply-Chain Security: Safetensors, Hashes, Licences

How to secure the open-weight model supply chain: safetensors, pinned hashes, signing, ML-BOMs and licence review.

Swfte Journal / Security Guide

When you download an open-weight model you are installing a few hundred gigabytes of third-party artifacts into your production environment: weight files, tokenizers, configuration, sometimes custom code, and a licence that governs what you may do with all of it. Software teams learned to treat dependencies as an attack surface; model files deserve the same discipline. This post covers the controls that matter, in the order that gives the most protection for the least effort, and is the supply-chain half of our model testing methodology.

OWASP's Top 10 for LLM Applications 2025 lists supply chain as its third risk, covering compromised models, datasets, libraries and hosting providers, including poisoned base models and adapters from public hubs. The framing is right: the risk is not only malware in a file, it is anything you did not produce and did not verify.

The threats, plainly

There are four distinct problems, and they need different controls:

  1. Code execution when the file is loaded. Some formats can run code. This is the classic malicious-model attack.
  2. Tampering. The file is not the one the publisher released, or not the one you tested. Weights can be altered to degrade accuracy or insert backdoored behavior without any code at all.
  3. Hidden code outside the weights. Custom modeling code in the repository, loaded when you allow remote code.
  4. Rights and obligations. The licence does not allow your use, or carries obligations you cannot meet.

Pickle: the format that can execute code

Many older PyTorch checkpoints (.pt, .pth, .bin) use Python's pickle serialization. Deserializing a pickle is not a passive read; it can run arbitrary code. Hugging Face's own announcement of the safetensors audit called pickle "inherently unsafe" and noted that a malicious file posing as a model could give an attacker control of a user's machine.

This is not hypothetical. In February 2025, ReversingLabs reported two malicious pickle-based models on Hugging Face that evaded the platform's scanner. The attackers compressed the PyTorch archive with 7z instead of ZIP, so loading with the default function failed and the scanner did not reach the payload, a reverse shell placed at the start of the pickle stream. Hugging Face removed the models within a day and improved its scanner to handle broken pickle files.

The lesson is not that the platform failed; it is that scanning is a deny-list and deny-lists have gaps. Picklescan, ModelScan and similar tools reduce risk, and researchers have documented bypasses. Use them as a filter. Do not use them as the control.

Safetensors: remove the code-execution path

Safetensors is a format that stores raw tensor data with a JSON header and no executable content. Hugging Face, with EleutherAI and Stability AI, commissioned an external audit by Trail of Bits, published in May 2023. The auditors found no critical security flaw leading to arbitrary code execution, along with smaller issues such as imprecision in the format specification and missing validation that allowed polyglot files, which were fixed. The post is careful to say that proving a library has no flaws is impossible.

What safetensors does and does not do:

  • It does remove the pickle code-execution path for the weight file. Policy: load safetensors only; treat any pickle-based file as untrusted code and refuse it, or convert it in an isolated sandbox.
  • It does not stop a tampered tensor. An attacker with write access to a repository can still change the numbers.
  • It does not cover code around the weights. Loading flags such as trust_remote_code=True execute code from the repository.

Treat remote code as a separate decision. Leave it off by default. If a model genuinely needs custom code, review that code, pin it to a commit and run it in an isolated environment, or prefer a model that loads with a standard architecture.

Pin the revision: a hash, not a branch

A branch name or tag is a pointer that can move. A full commit hash names exactly one set of files. With the Hugging Face tooling:

hf download <repo_id> --revision <full-40-char-commit-hash> --local-dir ./model

The documentation says to use the full-length hash rather than a short one. Record that hash in your configuration, and in vLLM and similar engines pass it explicitly when you load by repository ID rather than from a local directory (vLLM exposes a --revision option for this). For offline and air-gapped environments, download in a staging step, ship the cache, and set HF_HUB_OFFLINE=1 on the serving host so the client uses only the local copy and does not reach out.

Verify and keep a manifest

Pinning the revision tells you what you meant to fetch. A manifest tells you the bytes did not change on the way:

(cd ./model && find . -type f -print0 | sort -z | xargs -0 sha256sum) > model.sha256
(cd ./model && sha256sum -c ../model.sha256)

Generate it where you first verify the artifact, store it with the artifact, and check it at every boundary: after copying to your environment, before loading, and after any transfer into an isolated network. A hash proves the file is the one you tested. It does not prove the model is benign, which is the job of evaluation.

Signatures and provenance

A hash from the same place you downloaded the file only proves consistency. Signatures prove who published it. The OpenSSF Model Signing project (sigstore/model-signing) produces and verifies signatures for ML models in any format and size, with sigstore, self-signed certificates or key pairs, and the OMS specification defines a detached signature over weights, configs, tokenizers and datasets as one unit. Version 1.0 launched in April 2025, and OpenSSF reports adoption in model hubs such as NVIDIA's NGC and Google's Kaggle. If a publisher you depend on signs releases, verify the signature as part of intake. If they do not, record that gap and rely more heavily on pinning, hashing, scanning and evaluation.

For inventory, CycloneDX added ML-BOM support in version 1.5 (June 2023), with a component type for machine-learning models and a model-card reference. An ML-BOM record per deployed model (name, source, commit hash, licence, hashes, evaluation references, serving image) turns "what is running?" from an investigation into a query.

Adapters, quantizations and containers are supply chain too

The model is rarely a single file:

  • Adapters and LoRA weights from public hubs are third-party artifacts that change behavior. They get the same intake: pinned, hashed, safetensors, evaluated.
  • Quantized variants are produced by someone, with some toolchain. Prefer producing them yourself from a verified base, or treat a third-party quantization as a new candidate with its own intake.
  • Inference engines and containers. Pin versions and image digests, build from known bases, and scan images. An engine upgrade changes behavior as well as security.
  • Datasets for fine-tuning and evaluation have provenance and licence questions of their own, and poisoning risk (OWASP lists data and model poisoning as its fourth risk).

Licence review is part of the supply chain

A licence is a dependency with legal effect. Read the licence file and any acceptable use policy for the exact checkpoint, record the text, the date and the commit it applied to, and re-check when you upgrade. Patterns, as of this writing:

FamilyTypical patternCheck
QwenApache 2.0 on open-weight releasesConfirm weights are actually published; some flagships are API-only
GemmaApache 2.0 from Gemma 4; earlier generations used Google's own termsWhich generation you run
MistralApache 2.0 for most recent open-weight releases, including Mistral 3Specialist models can be non-commercial
DeepSeekMIT on R1 and recent releases; earlier V3 used its own model licencePer-release
KimiModified MIT on K2: model name shown in the UI above 100 million monthly active users or 20 million US dollars monthly revenueScale clause; later releases may differ
LlamaCustom community licence with a 700 million monthly-active-user clauseLlama 4 acceptable use policy withholds multimodal model rights from EU-domiciled people and EU-headquartered companies

Note that Apache 2.0 grants no trademark rights, that licences cover weights and not training data, and that a separate acceptable use policy can sit on top. This is orientation, not legal advice; counsel should read the text for anything that ships.

A minimal intake policy

If you do nothing else, adopt these seven rules for every model, adapter and quantization that enters your environment:

  1. Refuse pickle-based files; accept safetensors only.
  2. Keep trust_remote_code off; review and pin any code that must run.
  3. Pin the full 40-character commit hash.
  4. Compute and store a SHA-256 manifest; verify at every boundary.
  5. Verify a publisher signature where one exists; note where it does not.
  6. Record the licence text, acceptable use policy and the date read.
  7. Run the evaluation gate (capability, safety, languages, cost) before promotion, and again after any change.

Then record the outcome in an ML-BOM-style entry so that the answer to "what is running and why was it approved" is a lookup. In Swfte's loop these are the Pick and Harden steps (deploy models hub), and the record is part of the audit trail.

Related reading: how to evaluate an open-source LLM before production, deploying an open-source LLM in the EU, step by step, on-premise and air-gapped deployment, and open-weight vs closed models for regulated teams.

Keep the conversation practical.

Turn an idea into a working next step.

Discuss your use case
0
0
0
0

Enjoyed this article?

Get more insights on AI and enterprise automation delivered to your inbox.

Automate the response with SecOps Agents

Autonomous security orchestration: triage, investigation and containment, with a full audit trail.