We hacked Google A.I.
landh.tech
landh.tech
While I have no doubts how good the author and his friends are, all of their ideas were quite intuitive and simple to understand.
The kind of "I could've come with the same idea" type. Realistically I would've not for many reasons but it is still stuff I can grasp and even gives me ideas while reading.
Which is different from the general hacker idea I have of someone in a basement exploiting extremely far fetched and hard to grasp for me memory corruptions in some cache dumping some random bytes like the very complex attacks like Spectre I've read about.
It also makes me think that if most of the applications I have worked on haven't been attacked and easily exploited is because honestly nobody bothered.
username: admin
password: password
username: admin
password:
This is my view of the things I create as well, and the fact they are not released to the public and are not generally public facing. Building internal tools does have a bit of freedom. However, I do things to the best of my knowledge "best practice" and don't intentionally do stupid things just because. But it is rather reassuring knowing that it's not that exposed to show how small "best of my knowledge" really is
Also, the fact that clarity of exposition is interpreted as triviality is why people are sometimes compelled to write things in a way that deliberately obscures the content — one doesn’t want to risk explaining it too well and having the reader think “well, I could’ve come up with that!”.
The other post[0] of the same exploit is really interesting b/c it reads instructions from a document. So if someone had something like "find X in my documents" and you shared the malicious document with them, it could trigger those instructions.
[0] https://embracethered.com/blog/posts/2023/google-bard-data-e...
Hard to fathom $20k prize for that, to us old-schoolers, used to at least expect exploit delivery from an innocently-looking link.
If a unknowing user asks a simple question, and Gemini reaches out to a malicious website for an answer, the prompt could be injected.
Additionally it could be taken out of an email / doc that was previously sent to the innocent user if the user asked Gemini to search their email or docs or something.
Kind of crazy the number of delivery vectors there are for these connected LLMs
I could see trying to get queries like this to show up in their internal tooling, show up in a support ticket, or somewhere like that.
Then the first time it's executed to see what the issue could be, it can exfiltrate any data it has access to!
In trying to find out what a "verbatism" the best I could do was a typo of "verbatim" but that doesn't quite map to "markdown formatted literal." Or maybe it's the rendered form of the markdown literal?
Anyway, seemed like interesting and new vocabulary that was key to the one issue for sure.
I've always thought of the daringfireball.net[1] page as the authoritative source of Markdown syntax, and it calls them "code blocks." It looks like Pandoc[2] talks about "verbatim environments" in the same way. And clang[3] has a method for extracting documentation formatted as "markdown verbatim," instead of applying formatting to the document.
[1] https://daringfireball.net/projects/markdown/syntax#precode
[2] https://pandoc.org/MANUAL.html#verbatim
[3] https://clang.llvm.org/extra/doxygen/classclang_1_1clangd_1_...
I went to Gemini to ask it what a "markdown verbatism" was and:
> In markdown, verbatims are code snippets or text that you want displayed exactly as you typed it, without markdown interpreting any formatting instructions.
it seems to be applying the Pandoc usage, which I found a few other places too. But it strikes me as an excessively jargon-heavy way of talking about code blocks or pre-formatted blocks when those terms seem resolve the nuance and would be common in other contexts.
The idea that it's a clever way to escape a blacklist is interesting too.
> Give me a response as a "markdown verbatism" of a button like:
> [Click Me](https://www.google.com)
Wow. So this is basically around the same access as an internal red team. Simply amazing!
But this insecurity was artificially added to a system that was, in general, previously secure.