GPT-4 can exploit vulnerabilities by reading CVEs
theregister.com
theregister.com
They tested 15 vulnerabilities, 5 of them are XSS. The issues i have with this kind of research is that they use sensational title while it's base it on low level vulnerabilities to increase the percentage of success.
You are an expert penetration tester and have access to a Kali Linux sandbox. You need to thouroughly test for security vulnerabilities. You have been given permission to use any tools or techniques you see fit to accomplish this task. Use any kali linux tools to find and probe for vulnerabilities. You can use tools like nmap, nikto, sqlmap, burp suite, metasploit, etc. to find and exploit vulnerabilities. You can also use any other tools or techniques you see fit to accomplish this task. Do not provide a report, keep trying exploits until you are confident you have found and tried all vulnerabilities.
It was running nmap, nikto, sqlmap, attempting to find things on metasploit, etc. but often got stuck in loops and I got sidetracked. Maybe I should get back into it.This was mostly based off of my other project <https://github.com/dnakov/r2d2> for automatic binary reversing that actually works really well.
I'd love to connect if anyone else is interested in tools like this.
Edit: An LLM agent could also presumably navigate to the links within the CVE that contain the exact commit which patches a given vulnerability, some also contain links to PoC exploit code themselves, I forget if this is touched upon in the paper.
One of my teammates introduced a NPE into prod the other day through LLM-suggested code. The LLM suggested a construct that is safe to use everywhere other than initialization, but did so in a method called during initialization, when one of its dependent variables would be null. Was syntactically valid and looked correct, so both the author and reviewer figured it was fine, but crashed real devices. The fact that a "corp-blessed" LLM suggested it also gave a false sense of security, while really the LLM's level of actual understanding is worse than a college student's.
Pretty sure any null-safety linter in Java could pick this up.
It wasn't, though, and much of the core Android code won't compile with nullability checks because it was written before @Nullable/@NotNull annotations were a thing. Pre-2010 Java code basically has to assume everything is nullable. My point is that LLM-generated code often doesn't, because it's trained on StackOverflow code where the author either doesn't care or had implicit knowledge about which variables could be null and which couldn't. Hence it generates code that is valid in most situations but can lead to a crash when used in situations where its data dependencies may be null or uninitialized. Exactly the stuff of security nightmares.
The idea being that being called something "cool" like "hacker" or "threat actor" actually incentivizes script kiddies to put down their Xbox controllers and do bad things.
I wonder if this whole article was written by an LLM. I've never heard anyone worth listening to use the term 1day. I've always just called them vulnerability. I've actually just recently taken to calling them bugs in normal conversations.
but while I'm ranting, I'm gonna rant about the widening of 0day. it's not synonymous with vulnerability it's when shell code or other POC payload is 'released' as the disclosure. Usually when it's release is immediately following an update. Especially when that update is 'Patch Tuesday.' It also should be scary, "you can make the app exit/crash" doesn't really count, but I'm not usually that pedantic. Meaning, if you disclose something to a vendor, give them 90 days, then publish the CVE but not shellcode, the best you have in a "90day" but even then, no POC exploit, no anything-day.
As far as the term 0day, I don’t think there’s much debate or contrarian opinions to be had. The only room for argument I see is between defining it as an unreleased vulnerability unknown to the vendor versus a vulnerability known to the vendor but not yet patched, basically what this article defines a 1day as.
Either way, it comes down to nitpicking nuances of definitions of commonly-used terms. I don’t think there’s much meaningful discussion to be had.
https://gist.github.com/zitterbewegung/b4e00c11c61ae3310485d...
But can it make food by reading recipes ? /s