Edit: To highlight a specific problem here: a classic target for SSRF is the instance metadata IP address[1]. This IP address is not on the generated blacklist. Worse you've made it harder to detect this problem in the future.
I don't want to recommend a fix here; you're selling the fix. You should consider hiring a security expert to determine if LLM is really up for this task.
[1] https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/instance...
A few things we're doing to combat this: 1) We've given the entire corpus of CWE's to Corgea, and how to fix things safely. This we've found from our testing and users does a really good job. I personally QA a lot of results (in the thousands) and we've not seen that to be a common problem. 2) Corgea is designed to require two sets of eye balls. The first being the security engineer and the second being the developer that reviews the PR. We hope, that things should be caught. Additionally, we believe our fixes will be better than what a developer does. There are over 900 CWE's, and it's really hard for engineers to know how to fix every issue. Googling answers and asking ChatGPT can lead to them introducing issues. 3) We provide in the product AI generated explanations on the fix and how it was appropriate. This is to educate non-experts on the topic. 4) We already have checks in place to make sure things aren't misbehaving, but we're rolling out soon a more advanced fix checker to make sure we didn't introduce any new vulnerabilities. Based on our testing, and 5) Finally we QA a lot every week, and run reports on the areas we're good at or not so good at to help us iterate.
I really like the idea and I think that you're right about the goal being to fix bugs "better than the average engineer". I don't think you've reached that bar.
Well, at least they are showing a real demo and not some made up results.
I think that overall the idea has some potential, but not sure we are there yet.
For the first one the SAST scanner reports to us issues based on lines and issue type, so we generate fixes isolated for that issue. We do not generate fixes for other vulnerabilities in the same file for the same finding in the same because we want to have one fix to one finding. There might be another issue reported on another issue, and we plan on allowing people to group fixes in the same file together.
Not sure if I'm missing something on the shell=True. It's in the vulnerable code, which is why it changed it. You have to scroll to the right in the code viewer. https://github.com/RhinoSecurityLabs/cloudgoat/blob/8ed1cf0e...
Is there something I'm missing?
As for the second, There is no shell=True for me in the demo but it is present in the code you sent. So maybe it is just a bug in the presentation somewhere.
We'll also take a look at what's causing this. It might be a browser issue.
Looking forward to the archeological audits of LLM-developed apps x years from now that are a total mystery to the product owners…
As I highlighted in my post, LLM's generally are still not in a position to replace a developer for more complex tasks and refactoring. We're in the early days of the technology, but we are seeing extremely strong improvements in it over the last year. We on the team have QA'd thousands of results for public, and private repositories. The private ones are particularly interesting because the LLM's do not have that in their corpus, and have seen very strong fix results.
Most people just assume we're wrapping around an LLM, but there's a lot that goes underneath the hood that needs to happen to ensure that fixes are going to be secure and correct. Here are the standards we're setting for fix quality:
- The fix needs to be best-practice and complete. A partial security fix isn't a security fix. This is something we're constantly working on. - Supporting the widest coverage in CWE's.
- Not introducing any breaking changes in the rest of the code. - Understanding the language, the framework being used, and any specific packages. For example, fixing an CSRF issue in Django is different than Flask. Both are python frameworks but approach it differently. - Reusing existing packages correctly to improve security and if it does need to add a package does so in a standard way. - Placing imports in the correct part of the file. - Not using deprecated or risky packages. - Avoiding LLM hallucinations. - Ensuring syntax and formatting are correct. - Follow the coding and naming convention in the file being fixed. - Making sure fixes are consistent within the same issue type. - Explain the fix properly and clearly so that someone can understand it. - Avoiding assumptions that could cause problems. - Not removing any code that is not part of the issue.
Our goal is to get to 90% - 95% accuracy in fixes this year, and we're on a trajectory to do that. I will be the first to say 100% accuracy is impossible, and our goal is to get it right more times than engineers would.
We take fix quality and transparency extremely seriously. We'll be publishing a whitepaper showing the accuracy in results because it's the right thing to do. I hope this helps.
This fix is to demonstrate more sophisticated fixes, and it does require human input to determine the correct domain and IP config. We are introducing in the product for the ability for humans to add additional context pre-fix generation, provide feedback to generate a new fix after it's been generated and edit the proposed fix. Users have asked for these tools because of scenarios that require more insight.