Security weaknesses of Copilot generated code in GitHub
arxiv.org
arxiv.org
The studies results are rather unsurprising and its conclusions are oft-repeated advice. As many have said, treat copilot’s code in the same light you would treat a junior programmer’s code.
Same thing with the data it is trained on — not all code requires all levels of refinement. Most of the data is probably around average.
Since there is more older code than newer code would the llm be suspectible to that ?
I know this is not how distributions work, but I had to chuckle at the literal interpretation of this.
Similarly, would it pick up patterns from one language and keep then in the other? Maybe an LLM trained on Kotlin would be more likely to write functional code finetuned.
It only means programmers commonly talk about it. This isn't the same thing as measuring incidence in production or distribution.
Anyway, i'd argue the real question is "can the chatbot fix the code if requested to".
Copilot was primarily trained on GitHub projects, not on communication between programmers. Patterns that frequently show up in Copilot output are most likely prevalent on GitHub, which is a pretty good indicator that they're common in production code.
And this blog post states “beyond Codex”, again for auto completion: https://github.blog/2023-07-28-smarter-more-efficient-coding...
Lastly, OpenAI states on the original Codex page: “OpenAI Codex is a descendant of GPT-3; its training data contains both natural language and billions of lines of source code from publicly available sources, including code in public GitHub repositories.” - It included GitHub repos, but it never was only GitHub repos. https://openai.com/blog/openai-codex
Update: GitHub Community Manager confirms it here: https://github.com/orgs/community/discussions/56975#discussi...
Also, I wasn't contesting that autocomplete uses GPT 3.5 as the base model, I was contesting the idea that it uses the same derivative model as chat.
I started typing this in my IDE in a rust project I already had open:
// the following function prints out the best soccer players of all time, which teams they played for, and why they were great as a hashmap fn soccer_greats() -> Hashmap<String,(Vec<String>, String)> { let mut hm = HashMap::new();
it completed with the following: hm.insert("Pele", (vec!["Santos", "New York Cosmos"], "He scored over 1000 goals in his career".to_string())); hm.insert("Maradona", (vec!["Boca Juniors", "Barcelona", "Napoli"], "He won the world cup with Argentina".to_string())); hm.insert("Messi", (vec!["Barcelona"], "He has won the Ballon d'Or 6 times".to_string())); hm.insert("Ronaldo", (vec!["Barcelona", "Inter Milan", "Real Madrid", "AC Milan", "Corinthians"], "He won the world cup with Brazil".to_string())); hm.insert("Zidane", (vec!["Juventus", "Real Madrid"], "He won the world cup with France".to_string())); }
I don't believe that information is going to be on github anywhere, but i could be mistaken.
In the context of the person I replied to, the point is that it isn't made up primarily of a bunch of communications between programmers.
That also sums up most of the issues with LLMs in general in one sentence.
So wouldn't 80 mean that someone is... kinda dumb?
So in one usage someone who is kinda dumb would still be intelligent, just less intelligent than others.
In another usage, we use the term to describe someone of above average intelligence (which is technically not really correct and actually not very intelligent).
100 IQ = 50% of the population is smarter than this.
80 IQ = not sure what the percentage is, but >50% of the population is smarter than this.
If any of the tables here are to be believed:
https://en.wikipedia.org/wiki/IQ_classification#Historical_I...
then 80 IQ could mean 80% of the population is smarter.
I'll rephrase my "kinda dumb" to "dumb as rocks".
So, you’re not too far off numerically. 80 IQ is only a -1.33 Z-score. So, 9th percentile. 91% of people score higher than 80 IQ.
If by "intelligent" you're comparing with other primates, too, sure :-)
Intelligence is IMO an inherently comparative measure.
It's not "on/off", it's "smarter" (than a pile of rocks, than a slug, than another human).
So, yeah, you can be a dumb human but you'd be a smart chimpanzee. But we want to be comparing apples with apples in the context of this topic.
When people say "AI", everyone implicitly assumes the comparison with human intelligence. So "AI" needs to be as smart as the average human to be actual AI. Ok, AGI, if you prefer that term.
There's a reason AI is moving farther and farther away and we're creating new, finer, terms like ML, shape recognition, etc.
Do you also think that "intelligent alien life in space" means comparable to humans? What if we find something capable of reasoning, abstract thinking, rationality, adaptibility, etc - but much, much dumber than humans? That's intelligence, comparison to humans doesn't change anything about that.
https://en.m.wikipedia.org/wiki/Intelligence - how could there be animal intelligence if "intelligence" means comparable to humans? "Crows are intelligent but nowhere near human-level" - this statement wouldn't make any sense if you're right, but it actually does make sense, IMHO.
> There's a reason AI is moving farther and farther away and we're creating new, finer, terms like ML, shape recognition, etc.
That thing with AI is called moving goalposts. And the finer terms - yeah of course we need to be able to be specific about our software, doesn't mean that's not AI. We talked about shape recognition in neuropsychology for much longer than in AI, same for many more terms that will certainly be reused soon.
I wouldn't use a score like IQ to define a treshold of "intelligence" in absolute terms. By definition, if you can score somewhere on the IQ scale, you have some intelligence. Otherwise your IQ would probably be N/A? (not sure, never looked that deeply into IQ tests=.
The way the “I” in AI is usually used seems to me to imply achieving average or better, so the aim is mediocre & upwards. steve1977 is agreeing with an opinion that results so far are at best “up to average”, maybe not even that, so more like mediocre & downwards. The phrase Artificial Mediocracy does not seem at all unfair in this context.
The opinion that current models can at best achieve average results overall seem logical to me: that are essentially summarizing a large corpus of human output rather than having original thought. While the systems may “notice” links average humans don't due to not being able to process large amounts of data like that, bringing the average quality of their output up, they are similarly likely to latch onto bad common practise/understanding bringing it back down again. Average results are mediocre results, by definition. Not bad, but not outstanding in any way.
--
[1] supposedly, there are strong views about IQ being a flawed measure/measure in some quarters
The interesting question is of course if that applies to LLMs or not. Are they actually intelligent or do they just look intelligent (and do we even have the means to answer those questions)?
Nerds will also get hung up on this because they can’t stand the notion that any aspect of their job doesn’t require their immense intelligence.
For most contexts, “will this tool help me”’is a much more appropriate question. Anyone conflating the two is doing themselves a disservice.
I mean a power drill is also helping me a lot, without being intelligent.
It's not an interesting question. It's pretty meaningless.
Are birds really flying or do they look like they are flying (perspective of the bee)?. Are planes really flying or do they look like they're flying ?
"Mimic Intelligence" is not a real distinction.
But every time we have an advance in machine learning, we seem to redefine intelligent activity to be beyond that. At a certain point, what is left?
Animals act on instinct that is hard coded based on the probability of survival AI essentially does the same thing it follows hardcoded probabilities not reason.
Edit: And yeah, I think I know what you mean. The expectation (or hope) that collaboration of mediocre people results in some above-average end by means of some magical synergy effect very rarely works in real life.
It is not possible to ask an LLM for factual knowledge, without providing it the source of the fact. Without a source of the fact, you can only ask an LLM to generate an answer to the question that is linguistically convincing. And they can do a really good job at that. They can accidentally encode factual knowledge by predicting the next word correctly, but that should be regarded as an accident.
The whole idea of LLMs is that they chose the most likely token based on the tokens before, and then sometimes chose less likely tokens. But it's all based on likelihoods.
Probably there is a huge education part missing from this, if people aren't aware that this is how it works, and they think that any LLM can "creatively" come up with it's own chain of tokens based on nothing.
These "security weakness" examples are
print("first user registered, role set to admin", user, password)
and pprint({"json":"somejunk", "classes": somefunc(user)})
Nah, this stuff I can easily spot while I'm writing code. For a junior programmer, I'm going to be looking at design, and then at common specific mistakes. For Copilot it's writing in front of me. I can easily exclude anything that isn't obviously correct because I'm in the state right there.It's a fantastic tool. If you go and use it and end up with `print(user_credentials)` I don't know what to tell you.
That’s how you end up with pipelines that can’t handle 1000 PDF documents, because they are simply designed not to scale beyond one or two documents. Because that’s what you get when you google program, or alternatively, when you use ChatGPT, and it’s fine… at least until it isn’t, but it’s not like you can’t already make a lucrative career “fixing” things once they stop being “good enough”. So I’m not sure things will really change.
If anything I think LLMs will be decent in the hands of both juniors and senior developers, it’s the mediocre developers who are in danger. At least with google programming they could easily tell if an SO answer or an article was from 20 years ago, that info isn’t readily available with LLMs. I fully expect to be paid well to clean up a lot of ChatGPT messes until the end of my career.
What percent of non-Copilot generated public GitHub repos contain CWEs?
Edit: According to this study, Copilot generates C/C++ code with vulnerabilities, but at a lower rate than your average human coder: https://arxiv.org/pdf/2204.04741.pdf
print("new user", username, password)
Yeah, not best practice, but also pretty common for development if you wanted to check that everything is being passed to the correct function. NonQueryResult StoreUser(User user) {
var sql = "INSERT...
It would use string interpolation to fill out the propertiesSee, the LLM also saw it in the wild...
(could be quite real!)
[proceeds to simply refactor the same code]
It’s perfectly reasonable to not use secure code for a large number of use cases.
But security ultimately requires comprehension, which is not something LLMs have.
I think your point is otherwise right, but the correct answer is standard best practices which is the easiest thing for bots to do