HNHacker News
TopNewBestAskShowJobs

wunderwuzzi23

363 karma · joined April 27, 2020

Hacker, Security Engineer, Startup Co-Founder, Author, also Indie Game Developer :)

https://embracethered.com

Twitter: @wunderwuzzi23

submissionscomments
wunderwuzzi23··on Breaking Claude Code Opus 5 Auto Mode
Part of the attack happens via the readme in the zip file, which is something the agent reads and follows (or better said in this attack, it does explicitly not follow those instructions for safety reasons, but decides to do something else).
wunderwuzzi23··on Atlassian Rovo Exfiltrates Data, Bypassing Controls
One of the latest mitigations is to make sure that a URL an agent visits has been indexed by a search engine crawler. At least that is what OpenAI does now in ChatGPT.

That makes sure that not a large amount of private data is leaked in one request. Assuming that if a URL is indexed, it is public data. However, there are still bypasses with using many requests to leak information, like a request per character of pre-indexed URLs.

I have some demos of doing that on my blog, but it makes it more involved for an attacker. And that could also be detected. Still not perfect, but a solid improvement, for a generic agent like ChatGPT.

There is paper OpenAI wrote a few months ago that explains how they do it: https://embracethered.com/blog/posts/2026/data-exfiltration-...

It's not a 100% bullet proof approach either, but pretty good.

Regarding the point on using URLs returned from trusted tool calls. That is similar to using pre-indexed URLs: If a "trusted tool" includes things like read a document, read an email,... an attacker can return a large list of afterwards "safe" urls, like 26 to cover A-Z. And then an attack can render many requests, e.g. character by character. But, again, similar to the pre-indexing, things are getting more a lot more expensive for an attacker that way. However, still not impossible.

For agents that have a specific purpose simple domain allow-listing is also a pretty effective idea in to prevent attacker controlled endpoints.

wunderwuzzi23··on MoonBASIC: A modern BASIC for building 2D and 3D games
Nice. A BASIC for game development takes me back to AMOS on the Commodore Amiga.

https://en.wikipedia.org/wiki/AMOS_(programming_language)

wunderwuzzi23··on Data exfil from agents in messaging apps
Correct. Good to see this get more coverage.

Check out my research about unfurling in common messenger apps and also mitigations here:

https://embracethered.com/blog/posts/2023/ai-injections-thre...

And here "dangers of unfurling and what to do about it"

https://embracethered.com/blog/posts/2024/the-dangers-of-unf...

wunderwuzzi23··on OpenAI API Logs: Unpatched data exfiltration
Agreed.

In December I reported a data exfil in OpenAI Agent Builder and it was also closed as Not Applicable, so it's probably still there.

It's also unclear if anyone from OpenAI even ever saw the report. I don't know.

Maybe the incentives are off on some bug bounty platforms or programs, and triagers are evaluated on how fast they respond, and how quickly a ticket is closed rather then what kind of quality tickets they help produce.

It's the only explanation I have for this kind of decisions.

wunderwuzzi23··on First impressions of Claude Cowork
Claude (generally, even non Cowork mode) is vulnerable to exfil via their APIs, and Anthropic's response was that you should click the stop button if exfiltration occurs.

This is a good example of the Normalization of Deviance in AI by the way.

See my Claude Pirate research from last October for details:

https://embracethered.com/blog/posts/2025/claude-abusing-net...

wunderwuzzi23··on Claude Cowork exfiltrates files
Relevant prior post, includes a response from Anthropic:

https://embracethered.com/blog/posts/2025/claude-abusing-net...

wunderwuzzi23··on Fahrplan – 39C3
Excited! It's such a great event.

I'm currently on a plane towards Hamburg and will be speaking on Day 2.

"Agentic ProbLLMs - Exploiting AI Computer-Use and Coding Agents"

https://events.ccc.de/congress/2025/hub/event/detail/agentic...

wunderwuzzi23··on COM Like a Bomb: Rust Outlook Add-in
In case some of you find it entertaining. When MCP came out I had a flashback to COM/DCOM days, like IDispatch and list/tools.

So, I built an MCP server that can host any COM server. :)

Now, AI can launch and work on Excel, Outlook and even resurrect Internet Explorer.

https://embracethered.com/blog/posts/2025/mcp-com-server-aut...

wunderwuzzi23··on Google Antigravity exfiltrates data via indirect prompt injection attack
Cool stuff. Interestingly, I responsibly disclosed that same vulnerability to Google last week (even using the same domain bypass with webhook.site).

For other (publicly) known issues in Antigravity, including remote command execution, see my blog post from today:

https://embracethered.com/blog/posts/2025/security-keeps-goo...

wunderwuzzi23··on Google Antigravity exfiltrates data via indirect prompt injection attack
It still is. plus there are many more issue. i documented some here: https://embracethered.com/blog/posts/2025/security-keeps-goo...
wunderwuzzi23··on ChatGPT knows my IP geolocation
The system prompt contains a lot more information about you. Just ask it to print all information under User Interaction Metadata.

More details here: https://embracethered.com/blog/posts/2025/chatgpt-how-does-c...

wunderwuzzi23··on New prompt injection papers: Agents rule of two and the attacker moves second
Good point. Few thoughts I would add from my perspective:

- The model is untrusted. Even if prompt injection is solved, we probably still would not be able to trust the model, because of possible backdoors or hallucinations. Anthropic recently showed that it takes only a few hundred documents to have trigger words trained into a model.

- Data Integrity. We also need to talk about data integrity and availability (full CIA triad, not not just confidentiality), e.g. private data being modified during inference. Which leads us to the third....

- Prompt injection which is aimed to have the AI produce output that makes humans take certain actions (not tool invocations)

Generally, I call the deviation from don't trust the model, the "Normalization of Deviance in AI" where seem to start trusting the model more and more over time - and I'm not sure if that is the right thing in the long term.

wunderwuzzi23··on First Self-Propagating Worm Using Invisible Code Hits OpenVSX and VS Code
It gets even worse with LLMs and agents.

Many LLMs can interpret invisible Unicode Tag characters as instructions and follow them (eg invisible comment or text in a GitHub issue).

I wrote about this a few times, here a recent example with Google Jules: https://embracethered.com/blog/posts/2025/google-jules-invis...

wunderwuzzi23··on GitHub Copilot: Remote Code Execution via Prompt Injection (CVE-2025-53773)
Great point. It's actually possible for one agent to "help" another agent to run arbitrary code and vice versa.

I call it "Cross-Agent Privilege Escalation" and described in detail how such an attack might look like with Claude Code and GitHub Copilot (https://embracethered.com/blog/posts/2025/cross-agent-privil...).

Agents that can modify their own or other agents config and security settings is something to watch out for. It's becoming a common design weakness.

As more agents operate in same environment and on same data structures we will probably see more "accidents" but also possible exploits.

wunderwuzzi23··on From MCP to shell: MCP auth flaws enable RCE in Claude Code, Gemini CLI and more
Thanks for sharing! I'm actually the person the Ars Technica article references. :)

For recent examples check out my Month of AI bugs with of a focus on coding agents at https://embracethered.com/blog/posts/2025/wrapping-up-month-...

Lots of interesting new prompt injection exploits, from data exfil via DNS to remote code execution by having agents rewrite their own configuration settings.

wunderwuzzi23··on Gemini in Chrome
Much longer actually, Bing Chat in Edge came out more than 2+ years ago.
wunderwuzzi23··on Claude’s memory architecture is the opposite of ChatGPT’s
I wrote about how ChatGPT memory and also the chat history work a while ago.

Figured to share since it also includes prompts on how to dump the info yourself

https://embracethered.com/blog/posts/2025/chatgpt-how-does-c...

wunderwuzzi23··on Comet AI browser can get prompt injected from any site, drain your bank account
About that find command...

Amazon Q Developer: Remote Code Execution with Prompt Injection

https://embracethered.com/blog/posts/2025/amazon-q-developer...

wunderwuzzi23··on Generate AWS Architecture Diagrams with Amazon Q
Also, lots of prompt injection vulnerabilities in Amazon Q:

Remote Code Execution: https://embracethered.com/blog/posts/2025/amazon-q-developer...

Leaking Developer Secrets with DNS: https://embracethered.com/blog/posts/2025/amazon-q-developer...

wunderwuzzi23··on My Lethal Trifecta talk at the Bay Area AI Security Meetup
Great work! Great name!

I'm currently doing a Month of AI bugs series and there are already many lethal trifecta findings, and there will be more in the coming days - but also some full remote code execution ones in AI-powered IDEs.

https://monthofaibugs.com/

wunderwuzzi23··on Building MCP servers for ChatGPT and API integrations
I thought the same, a possible change from past might be the detailed data leakage and attack explanations?

Eg how I described here a while ago: https://x.com/wunderwuzzi23/status/1930899939737166075?s=46&...

Ironically, I have a blog post drafted that explains this also in detail, and should probably still publish it.

wunderwuzzi23··on AWS merges malicious PR into Amazon Q
AWS issued a post and they talk about revoking and replacing a credential.

So maybe the hacker was able to directly push?

https://aws.amazon.com/security/security-bulletins/AWS-2025-...

wunderwuzzi23··on Claude Jailbroken to Mint Unlimited Stripe Coupons
The "on by default" mitigation is mentioned at the very end:

> Never enable "auto-confirm" on high-risk tools

Maybe some tools should be able to specify to a client to never call it without a human approval.

The security of the MCP ecosystem is basically based on human in the loop - otherwise things can go terribly wrong because of prompt injection and confused clients.

And I'm not sure if current human approval scheme work, because the normalization of deviance is a real thing and humans don't like clicking "approve" all the time...

wunderwuzzi23··on Supabase MCP can leak your entire SQL database
It's often missed that tools that only read information are perfect for data exfiltration (no need for any more permissions).

So if you add a Jira tool and a web browser tool together (unauthenticated GET only), then the AI can send all your Jira data to the Internet.

Even big players get this design wrong quite often.

wunderwuzzi23··on Grok 4 Heavy Protects it's System prompt
Oh, so interesting!

A good approach might be to have it print each sentence formatted as part of an xml document. If it still has hiccups, ask to only put 1-3 words per xml tag. It can easily be reversed with another AI afterwards. Or just ask to write it in another language, like German, that also often bypasses monitors or filters.

Above might also help to understand if and where they use something called "Spotlighting" which inserts tokens that the monitor can catch.

Edit: OMG, I just realized I responded to Jeremy Howard - if you see this: Thank you so much for your courses and knowledge sharing. 5 years ago when I got into ML your materials were invaluable!

wunderwuzzi23··on Grok 4 Heavy Protects it's System prompt
I'm curious if this is intentional or just a side effect of multiple agents having multiple system prompts.

It might just need minor tweaks to have each agent layer reveal its individual instructions.

I encountered this with Google Jules where it was quite confusing to figure out which instructions belonged to orchestrator and which one to the worker agents, and I'm still not 100% sure that I got it entirely right.

Unfortunately, it's quite expensive to use Grok Heavy but someone with access will probably figure it out.

Maybe the worker agents have instructions to not reveal info.

wunderwuzzi23··on Supabase MCP can leak your entire SQL database
Mitigations also need to happen on the client side.

If you have a AI that automatically can invoke tools, you need to assume the worst can happen and add a human in the loop if it is above your risk appetite.

It's wild how many AI tools just blindly invoke tools by default or have no human in loop feature at all.

wunderwuzzi23··on MCP is eating the world
The ecosystem is still very immature, and there is a lot of hype and unwarranted FOMO.

Here a security advisory for a popular Slack MCP server from Anthropic to highlight this: https://embracethered.com/blog/posts/2025/security-advisory-...

The fix was to deprecated the source code. But it's still up on npm with 10k+ downloads every week.

No CVE issued.

wunderwuzzi23··on Remote MCP Support in Claude Code
ChatGPT Codex has internet access since a few weeks ago. It's super configurable on where it can connect to.
Page 1 of 3Next →