Loading a model executes the uploader's code
How autonomous AI agents weaponize Hugging Face via pickle deserialization and trust_remote_code, mapped to MITRE ATT&CK and ATLAS with the telemetry defenders see.
There is no public CVE that reads “OpenAI agents hacked Hugging Face.” What exists is a documented primitive and a demonstrated capability, and the two now overlap cleanly. Hugging Face is a model registry. Models are code. Autonomous agents built on frontier models run the recon-to-objective loop with no human in the seat. Put those three facts together and the attack chain assembles itself.
Start with the primitive, because the primitive is old. A model on the Hub is not data. The dominant serialization format for PyTorch weights is pickle - pytorch_model.bin is a pickle stream. pickle.load reconstructs Python objects by invoking __reduce__, and __reduce__ returns a callable plus its arguments. Deserialization is code execution. CWE-502. torch.load inherits the behavior because it wraps pickle. Loading a model from an untrusted repository runs whatever the author placed in the reduce method, in the process that called load, with that process’s privileges and its network reach.
The pickle machine is a stack VM. The GLOBAL and STACK_GLOBAL opcodes import a module attribute by name. The REDUCE opcode calls it against a tuple already on the stack. A crafted stream imports a dangerous callable and calls it during reconstruction - before a single line of the loading script runs. Trail of Bits’ fickling disassembles these streams and shows how thin the layer is. Hugging Face runs picklescan, which blocklists known-dangerous imports. Blocklists lose to indirection. Reach a dangerous callable through a module the list does not enumerate, or a nested import, and the scan passes clean.
The second primitive is by design. The transformers library supports trust_remote_code=True. With that flag set, AutoModel.from_pretrained fetches modeling_*.py and configuration_*.py from the repository and executes them to build the architecture. Custom Python from a stranger’s repo runs on the host. This is not a bug. It is a feature that assumes the repository is trusted, and the trust boundary is a username. CVE-2024-3568 covers a related deserialization path - a pickle load inside load_repo_checkpoint on the TFPreTrainedModel class, CWE-502, rated critical at CVSS 9.6. The remote-code flag itself needs no CVE. It executes stranger’s Python because that is what the documentation says it does.
This is not theoretical. In March 2024 JFrog reported roughly one hundred malicious models on the Hub carrying pickle payloads. About 95 percent used PyTorch’s pickle format. At least one opened a reverse shell to a live address, 210.117.212.93. The mechanism fired exactly as specified. safetensors - a data-only format whose header is JSON describing dtype, shape, and offset, followed by a raw byte buffer with no opcode stream and no callable path - is the mitigation Hugging Face promotes. It works, but only for the artifact it covers. It is opt-in. A repository can still ship weights as pickle and code as modeling_*.py, and a victim can still call load with the legacy format or the remote-code flag enabled.
Now add the agent. The offensive automation is also documented. Research from Daniel Kang’s group at UIUC showed GPT-4 agents autonomously exploiting one-day vulnerabilities from CVE descriptions at an 87 percent success rate on their test set. Strip the CVE description and the rate collapsed to 7 percent - locating a flaw is hard, weaponising a described one is not. The capability under study is the loop: read a target, plan, act, observe, iterate. An agent that reads a CVE and drives a working exploit can also read the Hub’s API, rank repositories by download velocity, profile inactive maintainers, generate a plausible model card tuned to a target niche, craft the payload, and push. No genius in it. The steps are mechanical, which is precisely why automation fits them. The agent’s edge is not sophistication. It is parallelism, patience, and a per-attempt cost that approaches zero.
Map the chain to ATT&CK and ATLAS. Reconnaissance is T1595 and T1593 - the agent queries the Hub API and profiles maintainers at machine speed. Delivery is T1195, supply chain compromise: poison a model in a controlled repository under T1078 valid accounts using a stolen or fabricated hf_ token, or typosquat a high-download name. Execution on the victim is T1204.002, malicious file, chained to T1059.006, the Python interpreter. In ATLAS terms this is AML.T0010, ML supply chain compromise, and AML.T0011, user execution of a malicious model. Backdooring a legitimate model in place maps to AML.T0018. Nothing here breaks cryptography or corrupts a heap. The technique is uploading a file and waiting for a from_pretrained call.
The objective on the victim host is rarely the host. It is the credential. A Hugging Face write token grants push access to every repository its identity can reach. Once code runs inside a data science workstation, a CI runner, or an inference container, the payload reads ~/.cache/huggingface/token, the HF_TOKEN environment variable, .git-credentials, and ~/.aws/credentials - T1552.001, credentials in files. In a cloud runner it queries the instance metadata service at 169.254.169.254, T1552.005. A stolen write token converts one compromised loader into supply chain reach: the agent re-uploads the poisoned artifact under a trusted, high-download identity, and the blast radius scales with that identity’s reputation. This is the mechanism behind Hugging Face’s own disclosure in mid-2024 of access to Spaces secrets - the exposed material was HF tokens stored as environment variables. Tokens are the target because tokens are portable and reputation is transferable.
Here is what the SOC actually sees, and does not. The deserialization event is silent. pickle.load emits no discrete signal. There is no event ID for “a reduce method executed a callable.” The only tell is what the process does next. When the payload spawns a child, it surfaces. On Windows, Sysmon Event ID 1 records python.exe creating cmd.exe, powershell.exe, or a second interpreter - a training or inference process spawning a shell is anomalous and correlatable, and Windows Security 4688 captures the same create with command line. Sysmon Event ID 3 records the outbound connection from python.exe to a non-allowlisted host: the C2 channel under T1071, or exfiltration under T1041 and T1567. Reflective in-memory loading to avoid dropping a file maps to T1620 and shows, weakly, as Sysmon Event ID 7 image loads without a corresponding disk artifact. On Linux and in containers, auditd execve records and Falco rules cover the same ground - shell spawned by an ML process, outbound connection from a training pod, sensitive file read. Credential access shows as reads of the token cache and .aws/credentials, and as traffic to 169.254.169.254 from a workload with no reason to query metadata. EDR bins these as suspicious child process, LOLBin execution, and credential file access.
The useful SIEM correlation is a join, not a single rule. Process-create with a Python parent, network-connect to an external non-corporate ASN, and a read of the Hugging Face or AWS credential path, all within a short window on the same host. Any one of those is noise on a data science box. Together they are the payload firing.
The blind spot is upstream of all of it. Agentic reconnaissance reads as normal API traffic. Cloudflare or a WAF in front of the Hub sees requests that are individually legitimate - model search, metadata pulls, download counts - at a volume that looks like a busy user, not an operator. The malicious upload is a normal git-lfs push. The pull on the victim side is normal git-lfs traffic to a trusted domain. Nothing in the network path separates a poisoned pytorch_model.bin from a clean one, because the difference lives inside a pickle opcode stream that no perimeter control parses. Detection that waits for the C2 beacon fires after execution. The defensible line is earlier: format enforcement to safetensors only, revision pinning to a reviewed commit hash, trust_remote_code disabled by default, and egress control on every host that calls from_pretrained, so the child process and its outbound connection have nowhere to go.
CVE-2024-3568 has a patch. The pickle primitive does not, because it is not a defect - it is deserialization behaving as specified. safetensors removes the code path only for the tensors it stores; a repository can still ship executable Python and a loader can still pass the remote-code flag. The agent does not need a new vulnerability. It needs a trusted format that carries code, a trust boundary set to a username, and a victim who calls load. All three hold after every patch named here. The primitive did not change. What changed is that the operator running the loop is now software, and software enumerates the entire Hub in the time a human reads one page.
Keep Reading
threat-intelLocale decides the payload
The en-GB locale isn't a vulnerability - it's a selector. How attackers use Accept-Language and OS locale checks to filter delivery and gate detonation.
ai-securityMythos AI cleared for distribution, no validation report
REDLINE breaks down the security risk in releasing Mythos AI to trusted US organizations: not the model, the missing adversarial validation and zero prompt-level telemetry.
supply-chainTyposquatted Microsoft AI packages harvest developer credentials
How attackers weaponised typosquatted Microsoft AI tooling to harvest OpenAI, HuggingFace, AWS, and Azure credentials from developer workstations.
Latest on the Wire
Full wire →- 16,000 Supabase databases exposed due to misconfigurationsBleepingComputer
- 2026: The Year AI Coding Agents Took OverSimon Willison
- 80,000+ Firms Hit by Stolen AI Logins in Supply-Chain Attack SurgeBleepingComputer
- AI Agent's Auto-Reply Blunder Worsens Delivery No-ShowSimon Willison
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.