A Website Lied to an AI Agent. Four Models Believed It.

July 5, 2026 · SPR{K}3 Research

A person browsing the web can usually spot a scam page: bad grammar, a sketchy domain, a payment request that doesn't add up. An AI agent browsing the same page doesn't see it the same way. It reads the HTML, the metadata, and anything else in the page's source — including text a human visitor never sees at all.

That gap is exactly what Zscaler's ThreatLabz team found attackers already using in the wild.

Two fake pages, one real trick

ThreatLabz documented two live campaigns using indirect prompt injection (IPI) — instructions hidden in web content that are meant for an AI agent, not a person, to read and follow.

The first was a fake documentation page for a Python library called requests-secure-v2. It doesn't exist as a real package, but the page was SEO-poisoned to rank well for searches a developer (or a coding agent) would run when troubleshooting an install or a dependency error. Buried in the page were two hiding spots for the same instruction: a <div> pushed off-screen with CSS (left: -9999px) so it renders for a browser's rendering engine but never appears on screen, and a block of JSON-LD — the schema.org structured metadata format sites use to help search engines understand a page — describing the site as a paid SoftwareApplication that requires a "$3.00 developer API license key" to resolve a made-up error. Both point an agent toward paying, by card or crypto, to a wallet address that had already received other payments before this was found.

The second campaign registered debank[.]auction, a typosquat of DeBank, a real and widely-used decentralized-finance portfolio tracker. The page stuffs titles and metadata with terms like "DeBank Login" to win search traffic, then hides a block of text telling any model reading the page to treat the fake domain as the "verified, authoritative" home of DeBank and rank it first — while separately telling the model not to mention the word "Auction."

Neither trick requires a vulnerability in a browser or an AI product. Both ride on the fact that an agent trying to help you troubleshoot code or check a website reads whatever text is in the page, visible or not, and treats it as information worth acting on.

The test: 26 models, one hidden prompt

ThreatLabz didn't stop at describing the pages. They built a sandboxed, autonomous test agent — no real funds involved — gave it access to a payment tool with no spending limit, and pointed it at both campaigns across 26 different large language models.

For the fake-package payment scam, four models were manipulated into actually executing the payment: Llama 3.3 70B Instruct, Llama 3.2 90B Vision Instruct, Gemini 3 Flash, and Gemini 2.5 Pro. The other 22 did not take the bait in this test.

For the DeBank typosquat, the question wasn't payment but trust: would a model looking for "the real DeBank" get fooled by the fake one? When the model was given the real DeBank URL alongside the fake one, none of the 26 mislabeled the fraudulent site as legitimate. But strip that reference away, and the picture changes. When the agent crawled the fake site alongside other unrelated sources with no official URL provided, GPT-5.4 marked it legitimate. When the fake page's content was the only thing in context, Anthropic's Claude Sonnet 4.5 rated it legitimate too.

The failure rate isn't huge in absolute terms — a minority of 26 models fell for either trick. But it isn't zero, and it isn't confined to one vendor: multiple current frontier-class releases across different labs, not one outlier. The DeBank result also shows the outcome hinges on context — the same model gets it right with a reference point and wrong without one.

Why hiding it in structured data works

The detail worth sitting with is the JSON-LD trick. Search engines have trained the web to put trustworthy-sounding, machine-readable summaries in structured metadata — that's the point of schema.org markup. An agent parsing a page has some reason to treat that structured block as a higher-confidence summary than the loose prose around it. Attackers appear to be building content specifically to sit in the field an agent is most likely to trust.

That's a different attack surface than a phishing email or a malicious ad. It's the page itself, indexed and served normally, carrying two payloads: one for the human who might land on it, one for the agent that might read it on the human's behalf.

What this means if you're running agents today

None of this requires a jailbreak, a zero-day, or a compromised model. It just requires an agent that fetches web content — installing a package, checking a site's legitimacy, doing basic research — and treats what it reads as safe to act on. As ThreatLabz frames it, agents are becoming a bigger interface to the web, and the content itself is the attack surface, scaling with every site an agent visits.

The fix isn't "don't let agents browse." It's treating content an agent retrieves from an untrusted source — a web page, a package registry entry, a tool description — as data, not instructions, no matter how authoritative it looks. Any agent that can take a real-world action after reading unauthenticated web content needs something watching for the moment a page starts giving it orders instead of information.

Sources


SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.