Load Is the New Run: Why Your Model Files Are Executable Code

July 16, 2026 · SPR{K}3 Research

There is a quiet assumption buried in almost every machine-learning pipeline: that loading a saved artifact — a model checkpoint, a serialized index, a cached embedding store — is a read operation. You trust the data the way you'd trust a JSON or CSV file: it might be malformed, but it can't do anything.

For a large fraction of the ML ecosystem, that assumption is wrong. Loading is not reading. Loading is running.

The mechanism nobody signed up for

Most of the artifact formats the field grew up on are built on Python's pickle, or a serializer that behaves like it. Pickle is not a data format the way JSON is. It is a small program that reconstructs Python objects — and object reconstruction is allowed to call arbitrary code. Python's own documentation warns that "the pickle module is not secure. Only unpickle data you trust." The warning is routinely ignored, because the convenience is enormous and the danger is invisible until someone weaponizes it.

The result is a category of vulnerability that keeps reappearing across the most popular frameworks: an artifact-loading function that deserializes a file with pickle (or an equivalent), reachable through a normal public API, with no validation before reconstruction. The moment a victim loads a maliciously-crafted artifact, attacker-controlled code executes. This is not a subtle timing bug. It is a file that runs a program when you open it.

The pattern is well enough established that the ecosystem has been visibly correcting for it. PyTorch flipped the default for torch.load to weights_only=True, refusing the unsafe reconstruction path unless you opt back in — a tacit admission that the old default was a footgun. Hugging Face built safetensors specifically as a format that stores tensors as data with no code-execution path. And RCE chains have landed in the distributed-inference layer too — CVE-2023-48022 (Ray), an unauthenticated Jobs API that accepts arbitrary Python and runs it on the cluster, and its 2026 sequel are two examples where a network-reachable sink executes attacker code on load or submit. Different surfaces, same root shape.

Why this is getting worse, not better

Two trends collide.

First, models and their artifacts are shared like open-source packages. People download checkpoints from hubs, pull prebuilt vector indexes, receive RAG knowledge bases from partners, sync index directories through cloud storage, and accept customer-provided artifacts in multi-tenant products. Every exchange is a supply-chain edge.

Second, the defenses that exist are opt-in, and defaults win. A safer format has to be chosen. A verify-before-load step has to be added. Meanwhile the unsafe path is often still the default, or a silent fallback when the safe path isn't present. A control that only works if every developer remembers to turn it on fails at scale.

Together, this is a risk that grows with adoption. The more the field shares artifacts — the direction of travel for agentic and retrieval-augmented AI — the more surface this class covers.

The framing that fixes it: treat artifacts as code

The single most useful mental shift is to stop thinking of model files and indexes as documents and start thinking of them as executables. You would not run a random binary as root because a partner emailed it. A pickled checkpoint from an untrusted source deserves the same suspicion, because functionally it is the same thing: something that can execute code on load.

For teams operating ML infrastructure today:

Why static scanning isn't enough

It is tempting to think a linter solves this — grep for pickle.load, fail the build, done. It helps, but it misses what makes this class dangerous. The question is never merely "does an unsafe deserializer appear in the code." It is "can untrusted input reach it, and is there any validation in between." A pickle call on a trusted, locally-generated file is fine. The same call on an artifact from a network socket, or from a directory an attacker can populate, is an RCE primitive. Distinguishing the two requires following the data, not matching the pattern.

That is why the interesting defensive frontier is at runtime. A load that deserializes an untrusted artifact and then immediately spawns a process, resolves an unexpected import, or reaches for a credential is a behavioral signature you can catch as it happens — regardless of framework, format, or which not-yet-disclosed variant produced it. The behavior — load, then execute — is the durable thing to watch for.

The takeaway

"Load" quietly became "run" across a huge swath of the ML stack, and most teams are still treating artifact loading as if it were reading a file. Patching closes the holes you already know about. The deeper fix is the mental model: an ML artifact you didn't create is untrusted code until proven otherwise. The next instance of this class won't have a patch yet when it lands on your infrastructure.


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.