Loading a trusted model runs untrusted code.
from_pretrained resolves a Hugging Face model name to whatever the branch holds now, extending a one-time trust decision over content it never re-checks.
A model reference on the Hugging Face Hub resolves to a set of files. When a process calls from_pretrained with a repository name, the transformers library contacts the Hub, retrieves the artifacts stored under that name, and loads them into memory. The system authenticates the location of the model. It does not, by default, re-establish that the contents behind that location are the same contents that were trusted when the reference was first written into the code.
This is documented behaviour, not a defect. The Hub is a distribution registry. It maps a human-readable string, org/model-name, onto versioned storage backed by git. A reference can point at a branch, a tag, or an immutable commit SHA. Most integrations point at a branch, usually main, and accept whatever that branch currently holds. The revision parameter exists precisely so that a caller can pin to a fixed commit, but pinning is optional, and the default resolves to the latest state of a pointer that is designed to move. The system is doing what a package registry does. It answers the question, what is at this name right now, and it answers it correctly every time.
So the frame here is not an intrusion. It is a resolution system operating exactly as specified. from_pretrained is a name lookup followed by a load. The load includes deserialization, and for checkpoints serialized with Python pickle, deserialization is execution. The safetensors format was introduced by Hugging Face to remove that property, to make loading weights a data operation rather than a code operation. But the loader still loads what the repository presents. Nothing in the resolution path asks whether the thing at the name today is the thing that was reviewed. The system was never built to ask that. It was built to fetch and to load.
The assumption underneath all of this is that a model, once evaluated and trusted, remains the same model. Trust was modeled as a property of the name rather than of the bytes. When a team vets org/model, benchmarks it, reads its card, and writes that string into a training pipeline or an inference service, the trust decision is made one time. Every subsequent pull inherits that decision without re-examining it. The name becomes a proxy for the artifact, and the proxy is treated as stable.
Inside that assumption are two smaller ones. The first is persistence: that the weights and configuration behind a reference at the moment of review will still be the weights and configuration behind it at the moment of use. The second is transferability: that trust granted to one state of the reference carries forward to later states, that version n+1 is trustworthy because version n was. Both are assumptions about time. Both treat a decision made in the past as valid in the present. The Hub does not encode either of these guarantees. A branch is a moving pointer by construction, and transferability across commits is something the consumer supplies, not something the registry enforces.
This is a reasonable model for a world where the entity that reviews the model and the entity that publishes it are the same, and where nothing between them can move. It is the same model that governs a trusted dependency in a language package index, where a version, once trusted, is assumed to remain trustworthy across future versions. Trust is granted to an identity and then allowed to persist and to transfer. In a static system, this is efficient and correct. It avoids re-validating something that has not changed. The efficiency depends entirely on the thing not changing.
What changed was not attacker capability, and not a mistake by anyone loading a model. What changed was the validity of the assumption itself. A git-backed reference is mutable. A branch pointer advances when a new commit lands. Weights can be replaced, a configuration file can be altered, a new file can be added to a repository, and the human-readable name over all of it stays identical. The reference is constant. The referent is not. The moment a commit changes what main points to, the state that was trusted and the state that will be loaded are no longer the same state.
The system does not detect this because it was never designed to. from_pretrained resolves the reference and loads the current contents, inheriting the trust decision that was made against a previous version of those contents. There is no revalidation step because trust was never stored as something tied to a specific set of bytes. It was stored implicitly, in a string, in a pipeline, in an assumption that the string still means what it meant. That assumption no longer holds once the underlying commit advances, and the system continues to behave as though it does. It resolves, it fetches, it loads, exactly as before, on contents it has never evaluated.
This is where the drift lives. Not in a single event, but in the widening gap between when trust was granted and what trust now covers. Over time a reference accumulates commits. Each pull silently extends a past decision over present content. The registry keeps its promise, which was only ever to return what is at the name now. The consumer keeps a different promise it made to itself, which was to trust what it reviewed. Those two promises were the same on the day the reference was chosen. They stop being the same the first time the pointer moves, and nothing in the system is watching the distance between them grow.
The failure is not a breach of the resolution path. It is the resolution path completing correctly. When from_pretrained is called with org/model and no revision argument, the transformers library asks the Hub for the current state of the default branch and receives it. Every hop in that exchange behaves as documented. The Hub returns the files under the name. The loader deserializes them. What is observable from outside is a successful load of the artifacts that exist at the name at that moment. There is no error, no warning, no divergence from the specified behaviour, because the specified behaviour was never to compare present bytes against past bytes.
The substitution is silent because identity of source stands in for integrity of content. The system confirms that the response came from the Hub, over TLS, under the requested name. It confirms the location. It does not confirm that the bytes at the location are the bytes that were trusted, because it holds no record of what those bytes were. The trust decision lived in a review that occurred once and in a string written into a pipeline. Neither of those is a checksum. from_pretrained cannot detect a change it has no baseline for. In the absence of a pinned commit SHA, there is nothing in the call that constitutes a claim about content at all. There is only a claim about a name.
This is where reference replaces validation. A commit SHA is content-addressed: it is a hash over the tree, and it can only ever refer to one state of the files. A branch name is not. main is a pointer, and a pointer is an indirection that resolves at read time to whatever it currently holds. When the consumer pins to a branch, they have written a reference that validates nothing about content and everything about naming. The safetensors format closes one specific hole in this, the one where deserialization of a pickle checkpoint is arbitrary code execution, by making the load a pure data operation. It removes the execution primitive. It does not restore the missing comparison. A safetensors file loaded from a moved pointer is still content that was never reviewed, loaded cleanly, exactly as designed. The format hardened the load. It did not give the reference a memory.
The pattern is execution on reference rather than verification of content. A system is handed a name that stands for a thing, it resolves the name at the moment of use, and it acts on whatever the name resolves to, without asking whether that thing is the thing the name was trusted to mean. The trust was granted to the name one time. The resolution happens every time. Between the grant and the resolution, the name is free to move, and the system carries the old trust forward across the gap without measuring it.
Take a container image tag. A deployment pipeline references an image as registry/app:latest. When it pulls, the registry resolves latest to a manifest, and the manifest resolves to a set of layers identified by their sha256 digests. The digest is content-addressed and immutable: it can only ever mean one image. The tag is not. latest is a mutable pointer, and it advances every time a new image is pushed under it. A pipeline pinned to the tag trusts a name that is designed to move. A pipeline pinned to the digest trusts the content directly. The mechanism is identical to the Hub. The trust decision was made against the image that latest meant on the day it was reviewed. Every subsequent pull inherits that decision over an image the reference has never re-examined. Same indirection, same silent extension of a past decision over present content, same registry keeping its only promise, which was to return what is at the tag now.
What both cases share is that the mutable reference is not a defect in the registry. It is a feature. A moving pointer is what lets a project ship an update without every consumer rewriting a hash. The convenience and the exposure are the same property viewed from two directions. The system optimized for distribution, for the ability to publish a new state and have it reach consumers automatically. Content-level trust would require the consumer to hold a baseline and compare against it on every load, which is precisely the re-validation the naming layer was built to avoid. In practice, the thing that makes the registry useful is the thing that makes inherited trust unsafe. A single unqualified reference cannot hold both the automatic propagation and the frozen guarantee. The reference resolves to one or the other, and by default it resolves to motion.
The system resolves the reference. It loads what it resolves. It does this the same way at review time and at every moment after, because it was built to answer one question, and it answers it correctly. Correctness is the problem. The load succeeds precisely because nothing in it was ever tasked with noticing that the content behind the name had changed.
Trust here was resolved a single time and then never again. It was fixed to a string and set loose over a pointer that moves. The distance between what was trusted and what is loaded is not an error state. It is the ordinary operating condition of a mutable reference under a static trust decision, and it widens on its own, quietly, with every commit that lands.
The control that reviewed the model still exists. The review happened. The benchmark ran. The card was read. None of it attached to the bytes, so none of it travels with them. The control exists. The outcome does not.
Keep Reading
supply-chain-securityYour build trusts whatever it can find
A build system executed code because a name resolved, not because content was verified. How trust attaches to the reference and outlives the artifact.
credential-exposureOne grep, full repo access
A security camera shipped a full-scope GitHub PAT in its login page bundle. The credential exposure, supply-chain exploit path, GitHub audit-log telemetry, and why rotation - not removal - is the only fix.
supply-chain-securityA valid go.sum hash proves nothing
Argegy is not a CVE. It's a Go supply chain claim against go-ethereum - module trust, init() execution, T1195, and where telemetry goes blind.
Latest on the Wire
Full wire →- 16,000 Supabase databases exposed due to misconfigurationsBleepingComputer
- 2026: The Year AI Coding Agents Took OverSimon Willison
- 80,000+ Firms Hit by Stolen AI Logins in Supply-Chain Attack SurgeBleepingComputer
- AI Agent's Auto-Reply Blunder Worsens Delivery No-ShowSimon Willison
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.