ChatGPT creates mostly insecure code, but won't tell you unless you ask
theregister.com
theregister.com
Much like with a lot of inexperienced, and unfortunately far too many experienced programmers, you'll get exactly what you ask for - nothing more, nothing less. So if you don't ask for security controls, you won't get security controls.
If you point out in your description what needs controls, you'll get a reasonable attempt at implementing what you've described. You still need to understand the output code enough to know whether you explained those controls clearly enough. Or just update the code manually yourself, akin to pair programming. That also works.
To be fair, this is a lot like reality. I have never seen a shop that doesn't produce insecure code unless you keep on top of them for proper implementation of security controls from the planning on up.
Exactly. The fact that lots of inexperienced programmers who lack the critical thinking and code review abilities will copy-paste what GPT is creating as is, is horrifying.
That's a strong take, and I know there will be very clever folks out there who make use of it who can apply common sense, but I think the majority of users have a strong Dunning-Kruger complex. I'm very unimpressed with ChatGPT as a tool for any kind of dev, and I largely consider myself to be a hack versus a lot of folks.
I've mentioned it before on HN, but with all the hype for ChatGPT, I wasted days trying to get it to spit out a correct NGINX config based on documentation fed to it. Hallucinated functionality that wasn't there, didn't implement URL rewrites correctly, and was just a colossal waste of time versus just doing it properly from scratch. Validating its output is much slower than just not using it and writing things manually IME.
Same goes for asking it for methods to export data from one program and import it into another, hallucinating all kinds of menu functionality that isn't there, commands that don't exist, and not considering file and format incompatibility.
Whatever people are using it for must be incredibly basic, but I don't know what could be more basic than an NGINX config. I can't even get sensible Home Assistant stuff out of it.
I can't believe people are asking ChatGPT anything instead of just reading the documentation when it can hallucinate some complete and utter garbage.
Would be nice if they interviewed people who knew the absolute basics of how this thing works before commenting on its properties
Inquisitive minds wish to know.
If you're going to claim it exists, do "the class" a solid by sharing. I, for one, am highly interested in anything that can reverse from blob -> potential source characterization.
SHAP is a great tool for unpacking deep nets on image recognition. Show which pixels mattered and which did not. Very cool.
Various dimension reduction or attention tracking or whatever exists for unpacking generate text and what not.
If you want to understand the nuance of a particular output you’re best off looking at the distribution of next tokens and their pre-temperature weighted sampling.
Different questions require different tools. Why did the model choose that phrasing? Why did the model choose this answer instead of that answer? What made the model make a certain claim?
A lot of it will be driven by the context assuming that the model works. So if you ask “Is X good or bad and why” the majority of text will be on the “why”, but the generation of that text will be contingent on the initial tokens that determined if it should say “good or bad” first. The model does not come up with a reasoning for its assertions, it comes up with an assertion and then creates a reasoning to defend it (in this question format at least). So perhaps that is all you care about?
Anyway, it depends on the question.
Breadcrumbs m'fellow. The quest for knowledge is made easier for everyone involved when we use nouns instead of assuming the other people in the room have access to the contents of our heads.
The "is it good and why" part we can read for ourselves. It's the name to search for to get started on the path that's important.
Apologies, though, t'was only meant to be a slight nudge. That post might have picked up a bit more acidity than warranted. Probably because missing antecedents are my #1 pet peeve.
I really think this sort of research is critical to do more of however as these outputs are going to be used in places where the assumption is that we're looking for these sorts of things.
As others have mentioned however, GPT generates insecure code, but so to most devs. The good thing though is that GPT can be systemically trained not to, while devs being more heterogeneous are more difficult to do the same. ;)
The main takeaway I have when thinking about tooling around GPT is that using it to generate things is fine assuming that you have a means to check the output against sensible criteria.
[1]: https://arxiv.org/abs/2304.09655
[2]: https://github.com/RaphaelKhoury/ProgramsGeneratedByChatGPT/...
The only difference now is the scale and ease to get your example.
The main problem I see with GPT-4 is that it does one small thing at a time. I have to manually drive it making the right requests.