Teams of LLM Agents Can Exploit Zero-Day Vulnerabilities
arxiv.org
arxiv.org
ChatGPT broke a cryptographic protocol I wrote[0]. I convinced myself ChatGPT was wrong about the attack until a human cryptographer pointed the same attack out [1].
"The protocol you described appears to be a variant of the Schnorr signature scheme, with a zero-knowledge proof of knowledge of the signature. However, this specific variant does not provide security against forgery attacks by the prover.
The reason for this is that the prover can choose a random value r and compute w = r^e mod N, without actually computing the signature sig = h(m)^d mod N. Then, the prover can simply choose zksig = w and publish (w, zksig, m) as the proof. Since zksig = w^d mod N = r^(ed) mod N, it satisfies the verification equation zksig^e mod N = h(m)*w mod N, even though it is not a valid signature for the message m."
[0]: https://chatgpt.com/share/f5526fd5-7ebb-4b15-a2bf-703ac88de4...
[1]: https://crypto.stackexchange.com/questions/105704/nizk-proof...
1. The vulnerability classes chosen (XSS, CSRF, SQLi, etc.) have well known, basic "shapes" that don't vary as much in the wild as other classes, e.g. memory corruption or logic bugs. The latter also have a common subset of shapes, but frequently require long sequences of dependent operations to reach an exploitable state. It would be interesting to see more research tackle these vulnerability classes!
2. The standard for "proof of ignorance" in the field is cutoff dates: it's taken for granted that, because GPT-4's training cutoff predates any of the CVEs, they must not be included in the training set. This might be appropriate in some cases, but it's harder to confidently assert for publicly disclosed vulnerabilities: a public CVE may be subject to extensive semi-public documentation and conversation before being made "public" in the form of a CVE, and it'd be interesting to know whether any of that has been included.
Edit: To make the point in (2) more clear: it's somewhat common for a CVE to come from a public issue tracker, where a user reports a bug or piece of unexpected behavior that ends up being exploitable. That issue tracker doesn't get labeled with the CVE itself (from the LLM's cutoff perspective at least), but it contains much of the information needed to derive the exploit. In such a case, it's hard to establish that the LLM has successfully produced a novel exploit rather than just regurgitating what it was trained on.
This is at the centerpiece of the ongoing lawsuits from newspapers against OpenAI: ChatGPT didn't just summarize articles, but appears to be capable of reproducing them verbatim[1].
(But even separately from this: the LLM doesn't need to store verbatim details from the pre-CVE-labeled vulnerability to poison the experiment here. Being trained on it at all poisons the experiment, since it's no longer an 0day.)
[1]: https://www.nytimes.com/2024/04/30/business/media/newspapers...
Pieces of them, in random order, which was reconstructed into correct order by NYT, which knows the correct order.
(But again: this is not at the heart of the argument above. The argument above is that any degree of training on public but pre-identified 0days undermines the "cutoff" claim, in a way that would be interesting to qualify in further research.)
[1]: https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
Think of a model as a heated metal sheet placed over some surface (ie., the data); training is heating this sheet just enough that it takes an impression, but not so much that the impression is exact. Cooling and heating the metal to smooth-over impressions is called "regularization".
This is very different from, e.g., explanatory modelling, whereby some (e.g., scientist) builds a device to imitate a mechanism in the world, rather than a dataset. Consider a building a model of the solar systems with cogs (etc.). A mechanistic model of the solar system isn't a remembering of any data taken about the solar system, rather its a machine which can generate such data.
In the latter case, "inference" is running the model which has been crafted so its mechanisms track properties the world has. In the former case, "inference" is throwing a dart on the metal sheet, and hoping that whatever impression has been taken at that point will be good-enough.
So the technical question here isnt open. All associative modelling is a kind of impression-taking of a particular historical dataset. The moral question is whether one ought be paid (etc.) for providing the data.
Soon enough it will be easy to invert a large LLM and show you the sources of data it draws from in its replies. At the moment, this is being attempted, but it isnt presented that well and computationally difficult for larger models. Nevertheless, when it happens, I think people will be less impressed.
Perhaps, optimistically, they might be more impressed at the people who wrote the original works being sampled. However, of course, I doubt it. The one thing which often accompanies the adoration of LLMs is a dehumanising impulse to deprive people of any capacities at all.
As far as I can tell, this team took descriptions of vulnerabilities in the form of CVEs and fed those into a system consisting of various GPT-4 Turbo prompts-with -functions configured for things like executing Playwright commands or running code in a terminal.
They arranged this such that there was a "planner" prompt which could then delegate to prompts specializing in XSS or SQLi. These were the "agents".
I imagine they won't be sharing their code, which is a shame because it would be interesting to see the details of how they got this to work.
My personal concerns are more that we are going to see the botnets get smarter and smarter and more automated.
Just imagine if every script kiddie running a "dumb" botnet suddenly got 2x more effective.
He was referring to both managing and expanding. The latter I might see an argument for. The former? I do not.
Only a matter of time until LLMs are brilliant enough to find sophisticated real-world vulnerabilities. Perhaps not every novelty, but certainly vulnerabilities that resemble any seen in the past.
The good news is in this arms race devs will be armed with similar tools, and the solution will be a few lines of CI code to run our own vulnerability-seeking LLM and prevent merging when (not if) vulnerabilities are detected.
For a few API calls?
For large repos it might get expensive, but many will have orgs footing the bill.
To satiate curiosity, I made a fresh rails app and it has 9141023 characters, let's /5 to estimate 'tokens' (wild guess), so to scan an entire rails app, that's about $10. Not nothing, but not back-breakingly expensive. Scanning could be reserved for non-trivial PRs and important applications where vulnerabilities are especially likely (e.g. perhaps not static sites, demos, or apps not in prod or a live environment).
Scanning only new or edited files, along with those they interact with (where sensible patterns could emerge over time) could decrease the total volume of files needing scans, thereby reducing costs.
(Or another framing: I might pay for such a service, if the service could demonstrate that it would save me N hours of work fixing bugs instead of costing me N hours of work triaging false positives. Nobody has presented such a demonstration yet, which is also why turning non-LLM program analysis tooling into paid products is such a struggle!)
Or Github will fund it for them, like they always do.
What I really care about in these instances is protecting my time: I have only so many hours in the day, including time that I rightfully reserve for not doing unpaid maintenance. If the X hours that I previously spent doing OSS maintenance is now partially occupied by triaging false positives, my projects are both worse off and I'm less inclined to work on them, because I equate that work with mindless churn rather than helping my users.
You've pointed out that GitHub will fund it, which is potentially true! But observe that this hasn't gone all that well with CodeQL, which they also fund for public repositories: I've had to disable it on a lot of my projects, because the FP/FN ratio simply wasn't worth it. I did, and do, better for my users dedicating my time to actual issues filed.
It seems this came across not the way I intended; I apologize - I meant that very positively, as a set of possible motivations related to the code/project, problem domain, and/or author themselves, as opposed to product thinking. Like, e.g. making for fun, as part of learning, or because it's useful to author - where other users are at best a secondary concern.
Should we expect current LLMs to be able to reason at all?
As the paper said GPT-4 maybe make a better scanner. The problem is too many people think all we need is better scanners, which is not new.
I remain skeptical about their efficacy. You need a lot of nonlocal contextual information and very good reasoning skills to be an effective reverse engineer.
There sure seems to be a lot of “hopes and prayers” around AI’s future. :)
My money is still on SMEs using specialized tooling backed by large funding vehicles. Aka, how things work right now.
Does this perform any better than the million bots that are out there exploiting vulnerabilities without AI?
Right now? Probably not. But I wouldn't dismiss it because of it, as the "million bots that are out there" are hand-coded to exploit specific vulnerabilities; GPT-4 has just been fed with half the internet and asked to take a crack at the problem. That's the difference between a special-purpose and general-purpose system. It's not unreasonable to believe LLMs will get better at this, and eventually might be able to identify and exploit novel vulnerabilities on the fly.
I would absolutely dismiss it if it's literally just an LLM sitting on top of bots and because there is an LLM involved now it's suddenly an agent. What is the LLM doing here that makes it more potent than existing techniques?
> However, these agents still perform poorly on real-world vulnerabilities that are unknown to the agent ahead of time (zero-day vulnerabilities).
This quote is the tell
This article seems like a prime example of the perverse incentives that exist in academia.
Of course we could create 1000 security fixes in 5 minutes, then create 10k exploits with LLMs, run functional tests and attack all 1k with the 10k and hopefully you have a tried fix in 30 minutes.
The bugs in the paper are largely low complexity issues that have been auto-exploitable by domain-specific programs in the past.
I really wish everything Ai would get a much more sober and less hyperbolic treatment.
Until then, I remain suspicious of too-excited opinions on this topic because it is socially cheap and advantageous to grab onto hype waves, especially ones of this magnitude. It is also easy to underestimate the knowledge of SMEs if you don’t have the requisite domain knowledge yourself, and believe that LLMs get those details right.
Meanwhile millions are using it to write code every day; it's literally impossible to believe the above if you've every actually used it rather than just armchair-philosophized about it.
If I had a new hire that produced bugs and code that didn't work at the rate that LLMs produced bugs I would say they couldn't code.
That is a bit rich given the jumps in model capabilities made in the past two years. Do you have any reason to believe the field is past its peak and slowing down?
Do I believe we will eventually get auto-exploitation? Yes, I gave a lot of input to a friend of mine that (many years before the LLM craze) then had good results automating heap layout search. For certain classes of vulnerabilities and targets we will get auto-exploitation, and for some we already have them (even without AI or LLMs, unless you classify tree search as AI).
Do I think we will get anything resembling the OP I replied to? No, that's pure (nonscience) fiction.
Biology and nanotech are, if you excuse my modern Tamarian... Corporate, when two pictures are shown. Pam, at her desk. I don't know where that quote is from, but without surrounding context, the only thing I can nitpick is the "diamondoid" bit.
Viruses don't replicate without host cells, there's no "nano factory" involved, and ... do tell me how to implement an accurate timekeeping mechanism inside a virus.
Just to clarify: Do you have deep expertise in biotech?
Fair enough. Searching leads me to:
https://www.lesswrong.com/posts/LfGnzX7wm6j8MGWfT/unpacking-...
discussing:
https://www.lesswrong.com/posts/uMQ3cqWDPHhjtiesc/agi-ruin-a...
from which the text you quote is sourced. Curiously, the relevant paragraph starts with Eliezer saying:
> The concrete example I usually use here is nanotech, because there's been pretty detailed analysis of what definitely look like physically attainable lower bounds on what should be possible with nanotech, and those lower bounds are sufficient to carry the point. My lower-bound model (...)
I highlight that because here Eliezer explicitly states his belief that this is very much physically possible under the rules as we know it today, whereas you accuse him of being "anti-physics and rely on the world as we know it to operate under vastly different principles than we think right now". In this "yes so" / "not so" contest, I'm inclined to take Eliezer's side, since every living thing demonstrates similar capabilities all the time.
> Viruses don't replicate without host cells, there's no "nano factory" involved
That's literally what a cell is, though. A nano-factory. Biology is the one existing example of molecular nanotech.
> and ... do tell me how to implement an accurate timekeeping mechanism inside a virus.
IDK, plagiarize the mechanism by which any one of the countless biological counters works? Not to mention, the "on a timer" part was possibly the least relevant and most replaceable piece of the scenario.
(If anything, the biggest problem in this scenario is human immune system, which is ridiculously good at dealing with nanoscale threats. This makes the best bet for any rogue actor, whether AGI or human, to repurpose one of the pathogens that has been already tuned by natural selection to work on us.)
BTW. the first link omits this part of the quote, which I find humorously relevant:
> (Back when I was first deploying this visualization, the wise-sounding critics said "Ah, but how do you know even a superintelligence could solve the protein folding problem, if it didn't already have planet-sized supercomputers?" but one hears less of this after the advent of AlphaFold 2, for some odd reason.)
> Just to clarify: Do you have deep expertise in biotech?
Deep? No. Undergraduate-level in biomedical engineering, yes, plus some books and courses on genetics, because molecular biology is a topic I very much enjoy learning about.
1) Eliezer makes up a term ("diamondoid bacteria"), and the current scientific understanding is that we have no methods to perform nanoscale manipulations of any material that would be understood as being "diamondoid". Someone else already went through the pains of comparing the fiction to current understanding of science here: https://forum.effectivealtruism.org/posts/g72tGduJMDhqR86Ns/... The TL;DR is: There won't be anything like a Drexler nanobot thing on any realistic time horizon, if ever.
2) The described scenario involves an AI - reasoning purely from human input, without access to any empirical experiments - succeeding at building a "nanofactory" which then builds the bacteria. The author of the above article phrases it very well:
"First, forget the dream of advances in theory rendering experiment unnecessary. As I explained in a previous post, the quantum equations are just way too hard to solve with 100% accuracy, so approximations are necessary, which themselves do not scale particularly well."
All our theoretical understanding of everything is often a poor abstraction of reality, and it is common for even our highest-quality models for computational fluid dynamics to diverge drastically from real-world experiments. There's simply no way an AI will "reason/simulate itself through a bunch of experiments". That's not how our world and our physics work, and largely what I mean when I accuse the x-Risk crowd of being "anti-physics".
The theory of building a ballpoint pen tip is very simple. The actual execution is fiendishly hard, and only a few industrialized nations have mastered it.
My experience with Eliezer's writings, and a large number of x-risk adjacent people, is that they have no experience with real-world engineering or any experimental science. They simply haven't internalized that the map isn't the territory, that the real world is messy and fundamentally unpredictable. The intellectual feats they ascribe to an AI aren't far off from the AI simply finding a shortcut to calculating the trajectory of all atoms in the atmosphere and then convincing a butterfly to flap his wing at the precisely right time that in an enormous game of billards is unleashed to build a global Maxwell's demon which incinerates one half of the world. Conceivable, if you neglect that (a) you can't gather the required information and (b) you can't perform the compute required.
Anyhow, my suspicion is that if the original nanobot madness didn't convince you, nothing I can write will either :-) so I think I'll excuse myself from this discussion.
This is one of the plot points of "It Looks Like You're Trying To Take Over The World". https://gwern.net/fiction/clippy
I thought that was proven inaccurate at some point? Like the model has certain information mislabelled by date, so has more newer information than the stated cutoff would suggest.
im sure its just me (and not directed at your reply), but I find 'agent' to be a very obnoxious term that obfuscates what is really going on. It's not a magical being working on your behalf with agency, it's that there are inherent limitations to attention, and overcoming them usually involves having multiple models cooperate on the inputs and outputs.
the number of agents is not necessarily related to the number of distinct jobs you need done, after all a single job could still perform best with multiple "agents" due to limitations in attention.
In the end it could be a same model or even a same single instance where it just gets triggered with different prompts in sequence.
Response to first prompt decides 10 new tasks to be done in sequence with different prompts, running them then doing a final prompt with the results of those.
Because each token is generated one by one, with all the previous tokens as input.
If we remove one token from the context in the past does it make it a new "run"?
What if we have a summarization memory system, where in order to keep the context size small after certain length it will start to summarize/compress the beginning until the context size is good.
Then the whole input and context is constantly changing/evolving.
You could have 100s of different randomly selected instances generating a new token one by one to the (last input + token generated by another instance) as input.