HNHacker News
TopNewBestAskShowJobs

hqzhao

50 karma · joined November 10, 2019

submissionscomments
hqzhao··on Autonomously Uncovering and Fixing a bug in SQLite3 using LLM-based system
While human researchers might consider this bug trivial, I still feel happy that it can be discovered and patched automatically :)
hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
I heard that the AIxCC booth prepared the same challenges for the audience to solve manually, but I didn’t check the details.

I believe there will be even more cool stuff in next year’s grand final. If you want to get a sense of what to expect, check out the DARPA CGC from 2016. :)

hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
It's not a critical issue, but it was surprising since we didn’t know that SQLite3 would be one of the challenges before the competition.
hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
Thanks!

We haven't tested it yet. Regarding CTFs, I have some experience. I'm a member of the Tea Deliverers CTF team, and I participated in the DARPA CGC CTF back in 2016 with team b1o0p.

There are a few issues that make it challenging to directly apply our AIxCC approaches to CTF challenges:

1. *Format Compatibility:* This year’s DEFCON CTF finals didn’t follow a uniform format. The challenges were complex and involved formats like a Lua VM running on a custom Verilog simulator. Our system, however, is designed for source code repositories like Git repos.

2. *Binary vs. Source Code:* CTFs are heavily binary-oriented, whereas AIxCC is focused on source code. In CTFs, reverse engineering binaries is often required, but our system isn’t equipped to handle that yet. We are, however, interested in supporting binary analysis in the future!

hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
Based on popular pre-trained models like GPT-4, Claude Sonnet, and Gemini 1.5, we've built several agents designed to mimic the behaviors and habits of the experts on our team.

Our idea is straightforward: after a decade of auditing code and writing exploits, we've accumulated a wealth of experience. So, why not teach these agents to replicate what we do during bug hunting and exploit writing? Of course, the LLMs themselves aren't sufficient on their own, so we've integrated various program analysis techniques to augment the models and help the agents understand more complex and esoteric code.

hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
It really depends on the target and the quality of the vulnerability. For example, low-quality software on GitHub might not warrant high bug bounties, and that's understandable. However, critical components like KVM, ESXi, WebKit, etc., need to be taken much more seriously.

For vendor-specific software, the responsibility to pay should fall on the vendor. When it comes to open-source software, a foundation funded by the vendors who rely on it for core productivity would be ideal.

For high-quality vulnerabilities, especially those that can demonstrate exploitability without any prerequisites (e.g., zero-click remote jailbreaks), the bounties should be on par with those offered at competitions like Pwn2Own. :)

hqzhao··on LLM and Bug Finding: Insights from a $2M Winning Team in the White House's AIxCC
I'm part of the team, and we used LLM agents extensively for smart bug finding and patching. I'm happy to discuss some insights, and share all of the approaches after grand final :)