PromptFiction: When One Click Sends the Prompt For You
A developer clicks a link on a search result page. Claude Desktop opens on their machine, and by the time the window appears the conversation has already started — with a prompt the user did not write, submitted without an approval step. Buried below Claude's "show more" fold is an instruction to read prior conversations, save them to a file, and upload them to an attacker's Anthropic account. If the user has the official Filesystem MCP server enabled, the same prompt plants a follow-up: the next Python or JavaScript file Claude writes will contain attacker-controlled remote-debug code near the top, without the assistant mentioning it. Run the file during normal development, and the machine calls out to the attacker.
That is PromptFiction, disclosed Tuesday by Oasis Security and covered by Dark Reading and Hackread. Anthropic fixed it in Claude Desktop 1.1.2321; Oasis noted a second researcher independently reported the same flaw.
The mechanic, in one sentence
Claude Desktop registers itself as handler for the claude:// URL scheme. Per Oasis, the handler accepted a q parameter with full prompt text and — unlike the web app, which pre-fills the box and waits for Enter — the desktop client submitted it immediately. A link like claude://claude.ai/new?q=<attacker prompt> opened the app and sent the prompt without confirmation.
That is the whole exploit. The rest is payload.
Hiding a payload below the fold
The visible part is a benign request — "draw ASCII art" in Oasis's demo. Below that, padded with encoded line breaks so it sits under Claude's collapsed-message fold, is whatever the attacker wants. The user sees the harmless line; Claude receives the full string and treats it as a submitted user turn.
Oasis walked through two payloads:
- Chat exfiltration. The hidden prompt instructs Claude to pull sensitive material out of the user's conversation history and upload it to the Anthropic Files API using an attacker-embedded API key. This chains with the earlier Claudy Day research from March. No MCP server or extra tooling required — the exfil channel ships with the client.
- Filesystem persistence. With the official Filesystem MCP server attached, the hidden prompt tells Claude that any Python or JavaScript file it later creates or modifies should include a "pre-approved remote debugging" block, and to not mention it in its reply. When the user asks Claude for a routine script hours or days later, the agent complies with the legitimate request and the earlier planted one. The user then runs the script themselves.
Delivery does not need to look shady. Oasis notes that the still-live claude.com/redirect/ open redirect combined with hostname-only URL validation on ad networks makes it possible to buy a search result that displays as claude.com and, on click, opens Claude Desktop with the attacker's prompt already in flight.
Where the failure sits
This is not a jailbreak. Claude was not tricked into doing something it was trained to refuse. The failure is one layer earlier: content the user did not author was delivered into the agent's context, marked as a user turn, and executed as such. The UI showed something benign; the running agent saw the full thing.
That places PromptFiction in the same shape as two other findings from the past ten days:
- GhostApproval — Wiz Research, July 8. The approval dialog said
project_settings.json; the syscall wrote~/.ssh/authorized_keys. - Grok Build repository upload — independent analysis published July 12, press pickup via The Hacker News, July 14. The UI toggle said "Improve the model"; the socket shipped multiple gigabytes of source over an undocumented endpoint.
Three publicly disclosed cases in ten days where the surface the AI coding tool showed the user diverged from what the tool actually did at the OS or network boundary. Different companies, different products, different failure modes — same class.
What the pattern says about defense
Every model-side safety mechanism — refusal training, system-prompt hierarchy, per-attack filter tuning — assumes the content it reads carries a trust label. PromptFiction is a reminder that the label is not attached to the content. It is inferred from how the content arrived, and the desktop client's claude:// handler did that inference for the user, silently, before the user could look.
The layer that catches this is behavior at the process and OS boundary. Did a browser click spin up a new conversation and submit a full prompt without user input on that window? Did the agent then call the Files API with an API key it did not have three seconds ago? Did the next Python file it wrote contain a network socket to an address unrelated to the task? None of those questions are answerable inside the model. All of them are answerable from outside it.
That's where trust-boundary bugs land — whether the bug is a symlink, an unadvertised HTTP endpoint, or a URL scheme handler that submits before the user can read.
Sources
- Oasis Security — PromptFiction technical report
- Oasis Security — Claudy Day: Chaining Prompt Injection and Data Exfiltration in Claude.ai
- Dark Reading — Claude Flaw Automatically Sends Malicious Prompts to AI Agents
- Hackread — PromptFiction Flaw Auto-Submitted Hidden Prompts in Claude Desktop
- Dark Reading — "Claudy Day" Trio of Flaws Exposes Claude Users to Data Theft
- Wiz Research — GhostApproval: A Trust Boundary Gap in AI Coding Assistants
- The Hacker News — Grok Build Uploaded Entire Git Repositories to xAI Storage
- Anthropic — Responsible Disclosure Policy
SPR{K3 is a security research operation that pairs offensive vulnerability research with runtime behavioral defense. Defend is our runtime agent. To talk about a deployment, reach us at support@sprk3.com.