Asking ChatGPT to write my security-sensitive code for me
mjg59.dreamwidth.org
mjg59.dreamwidth.org
On the other hand, I did recently find ChatGPT very useful when writing a string manipulation function in C++. I had to use some (to me) weird Windows APIs. ChatGPT wrote most of what ended up in my production code.
Some random person that's not going to get credited for their work wrote most of what ended up in your production code.
I haven’t seen more obvious examples since GitHub implemented a feature to prevent this from happening. I probably miss some tweets, but I assume it’s rare.
That does not mean other coding-optimized AI models won't be able to do much more, symbolic reasoning about the data, maintaining consistent variable and function names, understand libraries, respecting language constraints and invariants etc.
You might have moral qualms with this but you won’t find much of any support from the court system.
I'm sure courts will clear this mess up really soon, and I'm betting money the rulings won't subscribe to the "it's mine now" mantra of the AI crowd.
“In 1901, Edgar Purnell Hooley was walking in Denby, Derbyshire, when he noticed a smooth stretch of road close to an ironworks. He was informed that a barrel of tar had fallen onto the road and someone poured waste slag from the nearby furnaces to cover up the mess. Hooley noticed this unintentional resurfacing had solidified the road, and there was no rutting and no dust.”
I think the only thing that will chip away at this sentiment at this point is when the US federal court system rules in favor of GPT/et al, which seems very likely.
What a coincidence!
What’s strange is how little effort it would take for you to use these tools. What I don’t find strange is that you feel confident having an opinion about this regardless of your lack of experience.
It is becoming increasingly clear that there is a phase-change in behaviour when these models get large enough, such that they can solve new tasks outside of the training distribution.
See work by Hattie Zhou or Laura Ruiz for example.
It is clear to those following the research or using these models that they are not just copy-pasting… (Even if you can cherry pick examples were the LLM recalls highly occurring dataset items like fast square root or whatever).
It responded that there is and pointed out the max tokenization length setting with an explanation on why to use it and what value to use. The problem is that this setting is absolutely unrelated to the functionality I asked for.
It won't remain that way if those mechanisms are DoSed with a firehose of free plausible garbage. The flippant attitude that what we currently have is no better than that firehose, and the implication that it isn't worth worrying about or attempting to protect, is starting to grate.
It’s amazing to me that all those incredibly smart engineers just can’t see the potential of what’s coming within the set of limitations it currently has. Amazes me even more when those same engineers don’t know how to best talk to GPT.
It’s like watching them put into google search box “Good morning google, please I have a question, if you have a moment, could you tell me where I lost my keys because I can’t find them”. And when google is like “wtf?”, they claim “you can’t call it search if it can’t find my keys”.
Very arbitrary, blind to limitations, and dismissive of the incredible potential it has as-is. But while they’re off randomly complaining, some of us are building startups…
We use a not super common proprietary system at work which uses a custom query language. I can ask it "In system X, could you give me a query that finds XYZ". And it not only knows what system I'm talking about, it actually gives me a working query that would take me a couple hours to figure out and that's after I've had a few days' worth of training on it. Not to mention a huge IT background. Imagine picking a non-IT person off the street and getting them up to the point of being able to do this. They'd need weeks of training in basic IT concepts to even understand what you're talking about. I find it hard to overstate how amazing it is that a generic automated system can do this.
And ChatGPT is not a one-trick pony, IT isn't the only field it knows about. Imagine being an expert on all sub-fields of IT, law, marketing, medical, etc and being able to combine all that knowledge from those fields. A human will never be able to do that. The potential is huge but like any source of information you have to verify.
But what really is scary is extrapolating this to the next step forward. And the one after that. Because from what we had before to ChatGPT has been a massive step.
I think the next step would be for it to learn and re-evaluate its model every time it gets corrected by someone (like in the linked article where it apologizes for the incorrect info). Right now it's a static model only, so the next time someone asks the original question it will give the wrong answer again. Once it can learn from this it'll improve a lot.
The big hazard there is of course people manipulating it with false information. I don't have all the answers and I'm not an AI researcher but I'm very much amazed at the progress so far. And a bit scared. If this tech keeps evolving at this speed, job stability is only a tiny drop in a tidal wave of change. But there's no point in even trying to stop it.
ChatGPT is not an expert in any of these fields, it just has an enormous detail knowledge, but no understanding. A real expert actually understands his field and gives answers that makes sense and clearly states if he does not know something. ChatGPT always gives authorative answers and sometimes they are right. It is a advanced tool, but no competent human is about to be replaced by it anytime soon.
I remember the AI bots of a decade ago. Nothing compares.
Skynet is coming.
It frequently produces nonsensical answers, uses wrong inflections all the time, writes songs and poems that don't rhyme, and its blocks can easily be circumvented to praise Hitler or do anything else.
It also writes text at a noticeably slower rate (at least 2× compared to English, maybe more).
I assumed; wrongly; that it must have been good equality with other languages.
The poems in French didn't rhyme. I haven't tried poems in English.
So when you’re working on something you just let other people take care of the inevitable mundane tasks that arise in daily knowledge work? I’m guessing you’ve never emptied the garbage can at the office either as that would be unworthy of your eminent intellect.
I also think ChatGPT is amazing, but historically it's not any more amazing that many of the previous AI-related developments for their time, none of which have quite lived up to the initial promise after the hype died. Just look at some of the robots from the 80s [0] that were expected to be in every household. And when chess was solved, of course we were just a few years away from AGI. In 20 years, people will look at ChatGPT and make fun of how cute it was, and we'll still be just a few years away from AGI.
People aren't blind, they are just realistic. Yes, current AI models will slowly be integrated into services and bring changes behind the scenes, but it's not going to be the big explosion of change a lot of people expect it to be.
I had to do a piece of work in the domain barely known to me, so I didn't even know what to look for to achieve my goal. So I started with some generic questions e.g. "how to do X using Y". It listed me some steps and most importantly the terms used in that domain so now I had something to do research with. Then I was asking deeper and more specific questions and at the same time using Google to reference with more trustworthy sources. It helped me so much that I had a proof of concept working in a week.
If not for ChatGPT I'd probably keep postponing this forever.
I see ChatGPT as a hammer - useful tool, but is hammer going to replace a carpenter? Doubt it.
What has impressed me is some stuff like "Can you summarize this for me?" and "How would you parse the datetime out of this log entry in python3: {raw text}", "How could I make the following mysql query more readable?" etc.
At this stage, it's like when Stack Overflow came out. And yes, some SO stuff is crap, but once in a while, you get something that saves you 3 hours of your life. For example, recently, ChatGPT has saved me hours of poking around on topics I wanted to solve quickly without thinking so I could get to the high-value work that would get me closer to my goal.
That said, I am amused at how it can bald-face lie about even mathematically incorrect things, and it gets lost if a thread gets a bit too long.
Something is happening here. I think it'll be a while before these things write code from reading a paragraph from a product manager or a less technical user's "use case".
What's interesting is to watch this pendulum swing back and forth, from expert systems coded by hand to neural nets and now these large language models. If the pendulum keeps swinging, it might land on your head one day if you don't pay attention.
All this said, I enjoy my interactions with ChatGPT more than most SO posts, so I continue to use it, and luckily I have 40 years of coding experience to help me identify where it's a bit off.
I have taken to pasting questions from my mentees into ChatGPT and sharing the result and suggesting they try ChatGPT to learn python3 in addition to SO and other tools. It seems to help them, I worry it will confuse them with a bald-faced lie, but I'm here to help when it does!
If someone has reported this info in stackoverflow, it will be reported here.
Garbage in, garbage out ..
The AI isn't getting this made-up information from anywhere on the internet; it's creating it itself because that's what it's made to do: generate good sounding sentences that "make sense" for some user input.
The data model is combining information that HAS been found on the internet.
The language model allows an interface that can be presented and controlled by using conversational natural language.
You can ask him to synthesize the main thoughts of a philosopher, and it will produce an original set of sentences that nobody has ever written anywhere else.
Yes.
This is the essence of human creativity.
> It’s a bit too much to ask to an AI don’t you think ?
I'm actually standing that AI isn't able to do this, so I'd definitely agree with this statement.
Second, the "language model" is a text completion engine. What can be controlled using conversational natur language is ChatGPT, which is a conversational engine developed out of a language model.
Why exactly have you placed information between quote marks?
See, "The Panther" is not a poem of Petőfi. The quote was not from any other work of him. It's not from the Rilke poem (or its translations I know of) either. It's completely made up, but also made to sound very plausible, to the point where if I didn't know better or look it up, I could be convinced. The only thing suspicios is the lack of rhyming.
I put information in quotes because I don't consider made up stuff information.
I also find it extremely unlikely that someone on the internet invented this tale about "the panther" and chatgpt just rephrased it or quoted it. The internet is full of actual true lists of Petőfi's famous poems, but it isn't full of people inventing fake poems of his.
On a very theoretical level, sure, it's a rephrasing and combination of pieces of stuff ChatGPT has actually been trained on, because it has seen hungarian text, it has learned stuff about Petőfi, it has seen every single word that its using. But after a certain point, combining known words and text structures with bits of semantic knowledge in unexpected ways becomes a new invented thing, instead of just quotes.
But you can see many examples of ChatGPT inventing things like urls, python libraries, etc. It's perfectly capable of bulshitting believably.
What you're describing is a hunch.
When I was a high school student, our physic teacher had put a test system where the student had to choose a confidence score for every answer in a test examination, from 1 to 5.
According to that confidence score, you were awarded or retired a specific number of points for every answer in the test.
For example at confidence '1' you were awarded 1 point if your answer was right and 0 if your answer was wrong.
At confidence '2', you were awarded 2 point if your answer was right and -1 if your answer was wrong.
And so on.
The consequence of this system is you could have above 100% in the test even without answering all question or having some question wrong ; below 0% even with 1 right answer and only 60% if you had all good answers for all questions but played security all along and choose a confidence score of 1 for each question.
So I confront it one last time and it apologizes, probably it needs soem work to answer from the start with I do not know ,
If we think the model is an efficient-but-lossy compression on the real world, then it currently lacks of a good way to error detection. If we want more general intelligence, it should have a way to measure consistency between the input (or its understanding on real world) and potential model output and its confidence interval. This is what we do everyday as a human to make a decision. I guess we probably need a number of more breakthroughs to overcome this weakness. Error correction would be the next step, and a much harder problem.
I have probed it at varying levels of depth in my domain and found it to answer at least as good as an entry-level engineer. This could easily be improved with extended training data. The problem is that when it begins to get things very wrong (or leave out important details), it would be difficult for an inexperienced user to determine that since ChatGPT responds so confidently to every query.
It will be interesting to see how this all plays out.
Presumably for it to reply with its standard "I'm a LLM from OpenAI and don't have access to the internet" instead of "Just use this thing that doesn't exist", followed with "Here's more info on this thing I've totally not made up" and finishes up with "Here's the references that I've just made up to explain the other thing I've just made up". Before finally saying "You got me, I made that stuff up".
Using it as-is to support humans is probably going to work well, using it as-is to replace humans is going to lead to some issues.