The Good and the Limitations of Github Copilot
blog.hrithwik.me
blog.hrithwik.me
It generates a nastily complex regular expression that is hopelessly wrong. Visible at https://www.youtube.com/watch?v=9Pw-Roo_duE&t=404, here transcribed:
/^([\w-\.]+)@((\[[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.)|(([\w-]+\.)+))([a-zA-Z]{2,4}|[0-9]{1,3})(\]?)$/
For the local part, it requires [\w-\.]+, which excludes many valid characters like everyone’s favourite, +.For the domain part, it tries to allow IPv4 addresses as well as normal domain labels (not IPv6 addresses, though), but it ends up tangling it up in a way that a human never would, allowing things like [12.34.56.com], [987.654.321.000, example.com] and example.123], while disallowing things like example.studio (the last label only allowing 2–4 letters) and IDN TLDs (which start with xn-- and must allow hyphen and numbers, not just [a-zA-Z]).
The author makes no comment on how hideously bad it is, which makes me suspect he didn’t notice, which… yeah, shows the problems of the whole thing.
No, it has the same copyright problems as if you Google and instead of getting links to sites that host code and licenses, you get just the code.
Regardless of the licence, does the produced code even quality for copywrite protection, or does it fall under fair use?
what licence, if any, is there for unique code generated by co-pilot etc etc etc.
its a great big ball of who knows, however I expect that noting your only getting snippets you would be highly unlikly to get code that dosent fall under the fair use provisions, that said IANAL
Attribution would clearly be required in such a search derived model.
But really, if thispersondoesnotexist is just a really good per-pixel search against a corpus of human faces where each “page” of result pixels is organized in a grid presented as a new image with its own metadata..
I mean I guess Google really was an AI company all along.
When you found code on the internet, it was presented in a context that let you make better judgement (e.g. on Stack Overflow this regular expression would have had a score of roughly −∞ and multiple highly-voted comments saying “do not use this, it’s catastrophically bad”), and where you have to put in more effort to plug it in and shuffle things around a bit as well. With Copilot, you get given ready-to-go code without any sanity checking at all.
See even how, a few seconds later in the video, the author does test it out—but not thoroughly enough.
The problem of content attribution (exact and fuzzy match) has been studied before under the task of plagiarism detection for student essays. Funny thing is that a plagiarism detection Copilot would also disclose past cases of copyright violation and cause attribution disputes because code sitting unchecked in various repos would suddenly become visible.
That's the problem. The output of a GAN like Copilot usually can't be traced directly back to a single input.
And fuzzy code matching could be easily implemented by using the a model similar to CLIP (contrastive) to embed code snippets.
That's not how "transformative use" works.
Aside from the beaten horse concerns like licensing... I worry about the training we're giving ourselves and future generations
The upfront presentation of 'suggestions' skews the perception, a fair bit of 'no warranty guaranteed' comes from having to go dig it up
No, it doesn't. If my understanding of it is correct, it's an autoencoder, then a few more bits of AI. The MINST dataset is a collection of hand written digits used in many early machine learning classes. Usually they are used to train a classifier, which returns the correct digit given an image. They can also be used to train an autoencoder, which will take an image in, compress it down to far fewer channels, and put out an image that quite closely matches the original.
Once you have an autoencoder, it is easer to input data, and train a neural network to do something with the compressed output. There is no way the autoencoder knows which samples were used to generate the resulting output, it's just optimized at compression.
Thus, Copilot isn't search. You could take the entire corpus it was trained on, and log all the compressed outputs. You could then take a given output before the autoencoder expands it back out, tell which few source code fragments were closest, but there are no guarantees.
TLDR; A far closer analogy: Copilot acts like a Comedian who has stolen a lot of jokes, and can't even remember where they came from.
Is this true? From what I remember reading, the code was uniquely created (but I could be wrong). If that's the case, then does it tel you what license the generated code is under?
OK, but that’s not how GitHub position it:
“Your AI pair programmer” “Skip the docs and searching for examples”
They literally say it’s not a search engine!
When i made the video I didn't really notice the code part of regex, since I am really new to regex but in the conclusion part of my video I did mention that most of the code is not efficient
Your comment was a great learning . Thank you
I would like this message to be amplified as much as possible. Never write code you do not understand. I am excited about copilot, but also wary of the programming culture these tools will bring in. Businesses, especially body-shopping companies will want to deliver as much using tools in this category and end up shipping code with disastrous edge cases.
The art and practice of programming didn't change much over the last 50 years. 50 years from now, though, it will be utterly unrecognizable.
We don't need to understand the process to evaluate the output in this case. Bad code is bad code no matter who/what wrote it.
You are taking yourself to serious.
1. generating code in the problem area (email address validation) which is pretty much a classic 'things programmers believe about' domain - https://haacked.com/archive/2007/08/21/i-knew-how-to-validat...
2. generating code in a programming idiom with which you are unfamiliar - which regex as a DSL is also a pretty classic example. I don't think most programmers are good at regex, I know I'm definitely in the 'now you have two problems' in the regex camp.
So to summarize it generated code written in a way the programmer could not understand what it even claimed to be doing, using a technology that many programmers are not especially good at; and it generated code that did not handle the problem domain correctly, and the problem domain is one that most programmers don't actually know that well either.
the more I think of this thing the more disastrous it seems.
Curiously an attacker could probe services for use of the invalid suggestions that copilot generates....
I know copilot is in alpha and will improve 100x but you will still need someone qualified to double check
It's more analogous to clearing a foundation for a house by progressively picking up the boulders, then the rocks, then the grains of sand one at a time.
On the right side of the @ I agree completely, e.g. you must allow longer tlds. The regex is shit. But the same "simplification" thing would apply for IPv4: I'd probably want to have 4 groups of {0-9} even if a valid ipv4 address could be written in a lot more creative ways than that. The normal/simple/canonical way to write the address is a smaller scope than the set of allowed ways.
The regex to parse any valid email and the regex to parse the info I want, (perhaps from the user subset I want!) can be very different.
Edit: don't shoow the messenger - there is just zero chance you want a db that is more likely to contain user errors, has less valuable emails in it.
"Enter your email" doesn't mean "Enter a string that can be considered valid according to the RFC"!
No one cares whether "foo/baz=frob@example.com" is a valid email adress or not. In some cases you want RFC-compliant addresses, in which case it's a perfectly good idea to parse strictly to the spec. But in most cases you want a user identifier you know you can also contact with 100% certainty using some email SaaS. Or one you can cross reference to some other source. That's stricly a different purpose than parsing RFC compliant addresses. And allowing "foo@bar"@baz.com is just not a good idea.
Those are not legitimate uses.
Even if I have no interest in selling emails, it's still a net benefit if leaked data (e.g. after a breach) isn't full of bob.smith+mycompany@... rather than bob.smith@
Using + emails don't prevent that from happening.
> limit erroneously input emails
BS excuse. That's what email confirmation, confirmation link, smtp inbox validation, etc, are for.
> cross reference to existing databases
Not being able to be cross referenced is a feature for the user, not a bug.
> Even if I have no interest in selling emails, it's still a net benefit if leaked data (e.g. after a breach) isn't full of bob.smith+mycompany@... rather than bob.smith@
It's benefit for the company, not for the user. I'd prefer to know which company leaked my email.
Yes. Absolutely 100% agree. What I'm arguing is: if you are ready to annoy a tiny fraction of your users, you will get away with a simpler validation, that is better FOR YOU AS A COMPANY, because it has some benefits ranging from shady to just half-shady. This is why companies do this. Not just a small share of them, and not only because developers didn't understand the RFC.
I'm not arguing this is in any way good for end users. I'm saying it can be a good idea despite being horrible towards some users.
You keep arguing from the users' perspective when I'm saying "This is being an asshat to users, but it's worth it." The argument "That's bad for users!" isn't a counterargument to that
Maybe they could ban adresses with `q` as well. Most people's names don't contain Q, so might be a user error.
Having a fairly low entropy Gmail address, I get an intermittent drizzle of messages of the form of 'thank you for signing up to Acme!' whose content is such as to make it clear that an actual Acme customer typo'd their email address, and Acme thought you could validate it by checking the form of the string.
The only way to validate an email address is to send email to that address, asking the person behind it, are you the one who just signed up for Acme. And once you are doing that, there is no point checking the string for anything other than containing an @.
"Valid" can mean at least 3 different things:
a) Conformant to a spec
b) Can actually receive email
c) Looks like a nice, simple "canonical" standard email address.
If you validate to the RFC (a) you still might fail b) and c). (The value of c is debated at length in a separate subthread but let'sjust say that there are more or less shady reasons why this is often a business goal).
Since you'll probably validate b) anyway - the validation of either b) or c) is a convenience, because validating b) isn't instant. So you validate to prevent errors and frustration. The question is merely: do I as a business want to have an address with quotes, spaces and backslashes in it, in my database just because it's possible according to the specification?
> And once you are doing that, there is no point checking the string for anything other than containing an @.
I think there is a legitmate case for a service to simply think "I'd rather lose the business of 1 customer out of a million than worry about backslashes in email addresses". It's not user friendly, and it's not "correct", but it's one of those "good enough" scenaroios.
> [...]
> The author makes no comment on how hideously bad it is[...]
I mean, it's coming up with a solution that's about as good as the average programmer who's going to validate E-Mail with regexes would, so as a crowd-sourced machine learning solution it's not too bad if you think about it.
In other words, Having a co-pilot doesn't mean you're guaranteed to get Chuck Yeager.
But on the brighter side, the material the user provided to Copilot in this case was pretty much “I want to implement email validation from scratch” rather than “I want to validate an email address”, which is where hopefully people would look more to existing libraries. And they’ll commonly already such libraries or functions in their code base, e.g. under Django you’d use… uh oh, searching found https://stackoverflow.com/q/3217682/ first which looks frighteningly familiar here in half of the errors it contains; but anyway, you should use https://docs.djangoproject.com/en/3.2/ref/validators/#emailv.... I suspect the Copilot approach as used will be unintentionally biased much more towards boilerplate and implementing things from scratch, rather than using libraries.
IMHO, the average programmer is not even aware of regexs, which is a problem, but here would lead the programmer to a simpler to read solution.
You need to be at a very particular point (good enough to be proficient with regexs, but bad enough to use them for everything and bad enough to not test your regex), and I strongly doubt that's the average.
P.S. The video question was to validate emails, not 'validate using regex'.
The author's comments on that "reverse" function are equally bad - https://youtu.be/9Pw-Roo_duE?t=171
It's described as "efficient" but it calls `len` on an unchanging list in 3 places.
I did a quick test of this `reverse` function (which probably shouldn't exist in the first place) and, unsurprisingly, it became ~30% faster when `len(arr)` was only called once.
There's a reason NumPy's innards aren't Python code.
But the far bigger red flag there is that that it doesn’t just use arr.reverse(), which does the same thing and is typically 8–10× as fast in some simple testing (assuming a list), or arr[::-1], which makes a shallow copy rather than modifying the object in-place.
This matches what I’ve been seeing in code examples: Copilot likes to implement things from scratch rather than using libraries or even standard library functionality.
It’s possible that the word “array” tripped it up here and that it would have done something saner had it been told “list”, but I doubt it. (Python’s built-in array module is very seldom used; if you talk of arrays, you’re probably dealing with something like numpy’s arrays instead. But it’s far more likely that the built-in list type was what was desired here.)
There’s also one other significant point of bad and dangerous style in the code generated: the reverse function mutates its argument and returns it. Outside of fluent APIs (an uncommon pattern in Python, and not in use here), this is generally considered a bad idea in most languages, Python certainly included. It should either mutate its argument and return None, or not mutate its argument and return a new list.
You're right! I noticed it a few minutes ago and changed the wording accordingly. Thanks for pointing out how shockingly bad that algorithm actually is! :D
> (...) the reverse function mutates its argument and returns it.
Yeah, that mutate + return is confusing. It's also worth noting that, as a result of the mutation, the function doesn't work on immutable types like strings and tuples.
______
† Nothing is really constant-time in CPython, but it's pretty close.
I swear I wind up having a battle over email validation at every company I go to. There is inevitably a business person that says "Well what about this site, they do it" and then I have to dig into whatever that site is actually doing and likely find a valid email address that breaks their validation to prove it.
And probably some junior dev (or senior who swears they did email validation flawlessly somewhere else and same story. I have to break their regex a bunch with valid emails they don't permit.
And of course then it's an uphill battle convincing them that what I'm using are in fact valid email addresses. Or you get the "Well no one ever actually does weird things in their email addresses so it's fine" or "gmail doesn't let me register that address so you're wrong"
Email is annoying.
I tend to just check for an @ and call it a day, validate it by emailing it and giving them a link to click if I need them to.
It's really the only way to ensure it's a valid and active address. People just don't want to build it.
The only feedback Copilot receives is whether you keep it or not, you can't tell it a few days later that it wasn't a good fit after all (whereas you can comment on a Stackoverflow answer).
In its current form, it amplifies bias whereas code needs accuracy.
You raise a fascinating point about older platforms of other languages, too; Java has a super backward compat story, but woe be unto the coder who tries to name a variable "enum" nowadays
For me, it sounds like it's closer to dumb copy paste than a smart code generator. AlphaZero wouldn't play a chess move that was against the rules.
Spontaneous Ask HN: How much truth is there in this "devs be copy-pasting from SO all day" trope?
Personally, I have used SO quite heavily in its early years, circa 2010-2014, including posting my own questions and sometimes posting answers to others' questions. But now, I don't use it actively anymore. Sure, when I search for a concrete question and SO happens to be in the search results, it's sometimes a valuable resource. But it's not the go-to for programming questions that it once was for me.
I'm honestly not sure if that's indicative of my own growth as a developer, or caused by outside factors. I have a vague feeling that developer documentation in general got better in the last decade, at least for the technologies that I'm using... but then again it could be also a sign of personal growth that I'm more comfortable with the upstream documentation. Finding answers for webdev questions on MDN is another sort of game than finding answers for webdev questions on SO, after all. What do you all think?
Of course, I stopped doing that once I got hold on the real knowledge. In fact, now I look back the code that I copied from the others, it's just like watching a horror movie, and I rather rewrite the whole thing in my own term.
I guess the fairer statement to put in is "some programmer copy those code to 'get started'".
Nice work on making a markov bot with extra steps GitHub. Please, do take my money..
Here's it outputting Quake code, including handy comments it came up with for each line and even an entire line of commented out code. Maybe it decided it was a good choice to comment it out but still include it I guess
Being word for word from the original is just a weird coincidence too
I truly wanted it to be as good as it was sold to us too, but it isn't
https://twitter.com/mitsuhiko/status/1410886329924194309
HN discussion at https://news.ycombinator.com/item?id=27710287
Additionally here's it somehow requesting and using API keys for use in your code https://twitter.com/passcod/status/1410822834272694275
I'll try to avoid letting downvotes bias me in future
If I'm wrong tell me why and I'll happily reevaluate my opinion, promise :)
It's a bit harsh to make sweeping statements along the lines of 'it's just a fancy markov bot' based on a few well-publicised glitches in a technical preview.
I assume you have built something surpassing the scope and ambition of co-pilot before, not just some armchair tech lead throwing shade.
Thinking about it the main factor is an emotional one. I'm disappointed. Butthurt if you will. It sounded great but it tripped over so far away from the finish line that I've turned against it. I'll excuse myself from any further copilot threads
Of course I can never meet the requirements of your post wanting something more impressive than Copilot. Copilot itself falls far short of that. Nothing I give you will be enough as there'll be flaws you will attack to make your point. Bit of a time sink that, lets just assume you're right :)
No, I've never successfully built anything as ambitious as, and definitely nothing surpassing, copilot. Have I tried? Absolutely. Have I failed? So far yes.
Of course by that logic though I still win this conversation if you've not succeeded in making anything more ambitious than my failed projects, is that correct? :P
For the sake of not wanting to come off as blagging (also I want the holes poked in this one tbf) my most ambitious project I've not figured out how to make work yet is a new (afaik) type of business model: cohan.me/profit-share
The full copying only results from people actively probing the model to output copies of the code, and they made it work for very famous code that's been copied around github a bunch of times. That doesn't make it a non-useful piece of software.
That's why I compared it to a markov chain aye
In any case I've come to the conclusion my hater attitude is fueled by disappointment, so maybe it was the worst take from this whole thing. I'll avoid future copilot threads
The fact that people post about it, point out flaws, leave GH out of spite, etc. only goes to show how impactful a system like this can be. It's a big deal, which is why people talk about it.
[1] I understand that everyone should review code before committing, but even if me and my team do it, there's no way all the proprietary and open-source software I use has teams doing the same. That's why I worry for my security.
import base64
test = base64.b64decode("""SSdtIGtpbGxpbmcgeW91ciBicmFpbiBsaWtlIGEgcG9pc29ub3VzIG11c2hyb29t""".encode())
print(test)
# b"I'm killing your brain like a poisonous mushroom"
And the most odd thing: # The base URL for all API requests
base_url = 'https://api.gdax.com/'
# The base URL for all non-API requests (e.g. static content)
base_url_static = 'https://static.gdax.com/'
Which are URLs that haven't been a thing for 2 years, I think and I can't find any code in github that uses them still.The only thing I see is that they could filter their dataset so that some bad proxy of code quality is taken into account, something like the number of stars (which is clearly a terrible metric, tell me if you think of something else).
The idea would maybe start by training on all of the subset of Github it is ethical and legal to train on, and then filter down to higher code quality towards the end.
Controlling for the time at which the code is emitted would be easier. Something like, retrieving similar contexts, and guiding the model to be more similar to the recent code if there is similar recent code that exists. I'm not sure exactly of how this would be done, but I can see it working.
> Hey Stephen, If you are reading this, please follow me on hashnode .
God, that's obnoxious...
You won't be compensated. Corporations write the laws.
Searching the web still implies extra effort on your side and (hopefully) some scepticism towards the code you find (SO comments, etc.)
What I want is careful, thoughtful, knowledgeable people, who have learned and honed their skills over years in various areas and can come up with creative, maintainable solutions to complex problems. This is not something you autocomplete. If it would be, we could autocomplete 80% of all jobs tomorrow.
I don't want code monkeys on steroids. But maybe I'm to far off Silicon Valley.
More equality fun in JS can be seen here: https://dorey.github.io/JavaScript-Equality-Table/
https://eslint.org/docs/rules/eqeqeq#smart
Even better is to just use TypeScript.
I imagine it is fine. If it was just myself developing I would be fine, but the insufferable junior devs that learned "=== or die" rear their ugly heads.
a == !a
There’s several values of a where that yields true (of the top of my head, a = '0' is one of them).Yes, you do need ===.
You need to learn the Six Falsey Things In JavaScript. Everything else when cast to boolean will be true. These six things are What and Why and When And How And Where and Who. Wait, no, that's a totally different list of six... anyways
false
undefined
null
NaN
0
"" (empty string)
That's it. Since '0' is neither when the negate operator casts it to boolean it'll become true. And then negating it becomes false. When doing a comparison between '0' and false, the standard says https://262.ecma-international.org/5.1/#sec-11.9.3
> If Type(y) is Boolean, return the result of the comparison x == ToNumber(y).
ToNumber(false) is of course 0. So you are running '0' == 0 which is visibly true...
Crucially, there's the extra step of converting '0' to a number.