From what I understand, it is not proven that the AI uses the knowledge of concepts and logic to write the new code. It is likely that it actually performs instead a very optimized stitching of code it previously saw.
Is my understanding outdated here?
From the ethical point of view, I'd say you're making some assumptions here that result in it being ethical when a human does it, and those assumptions might not hold for an AI.
For example, you're assuming a win/win outcome, where your learnings from copyrighted open source code don't harm the original authors ability to find work, or the value of their code.
With an AI I think there are possibilities we're looking at a win/lose situation, where Microsoft wins big, and maybe some other developers that also profit of their use of copilot, but where the original authors of the code that went to train it see their skills be devalued over time as a direct consequence.
In my opinion, a win/lose is unethical. What I'm not convinced is that we're looking at a win/lose, but I think there's a possibility.
I have seen this line repeated many times, but I never saw it actually explained. A lookup table is dumb and easy to understand/interpret. Deep models are not that. They are also not a linear interpolation of … something. What exactly is the claim being made here? Yes, deep models don’t generalize too well on ood data. How does this make them a “very optimized” lookup table?
Here's the real problem:
We're on a forum in which most of the participants are in at least the top 0.1% of technical ability, and yet here we are waving our hands and speculating on "what AIs think" and how "they probably/likely compute" things.
Last week I met with a director of a new "AI" research group, chuffed to the nines with a massive research grant he just landed. I'm happy for him, with only one little concern - that he knows nothing whatsoever about machine learning or mathematics and outspokenly doesn't believe that "knowledge is that important in the new reality".
Copyright infringement is an inconvenience for people who worry about that sort of stuff. Sure. But don't you all see a much more serious issue? It's bad enough that code is already so precariously bloated and over-complex nobody bothers to debug critical applications any more. Now we want to add "assistance" from tools that nobody understands.
People have special rights and responsibilities under our legal system - we don't send an airplane that has a mechanical fault to prison for a crash nor do we extend right to life to a web server. Humans have an implicit the right to learn by reading copyrighted material, machines have no such right.
What if instead of Copilot, it was a bunch of humans who were searching all the source code they could access and then copying/autocompleting that code, regardless of the license.
Is that still OK? If yes, why?
Copying code, even if its from a mix of many different places, and the results look like a mosaic, would still be copying.
If the Mturk worker just suggested an implementation they came up with that be fine.
You’ve been authorised to see this code.
Nice you got something out of it, I'm not judging you either, but it was probably not the correct way to operate. What you should've done was notify the owners of the incorrectly configured system and left it at that.
You're also not a massive international conglomerate who should know better than to read every ones code and use it to turn a profit without first asking for permission.
I use Github like a bank, not a public library (unless I'm working on open source). I never would've allowed them to read through all my code and use it for profits without at least asking.
Illegal would imply some kind of intent or malice. I was legitimately trying to access the executed result, which I would have been authorized to do if the service was operating normally.
> What you should've done was notify the owners of the incorrectly configured system and left it at that.
Seems unrealistic to "leave it at that". I had to read the output to understand it wasn't what I expected, and once I read it I knew how to code, at least to a cursory degree. The code was simple and it was a service I used frequently, so it was immediately clear how the code translated to the results I was accustomed to. Maybe that would be harder to do that now in my old age, but I was just a kid so I had neural plasticity on my side.
> I use Github like a bank, not a public library (unless I'm working on open source). I never would've allowed them to read through all my code and use it for profits without at least asking.
I don't know what kind of banks you deal with, but banks normally do read through your banking records and use that information to sell services to their clients – notably loans, which require knowledge of your deposits to offer.
Misconfigured CGI handlers in Apache were very common in the late 90s, treating Perl as text/plain. There's no laws being broken, just a bad httpd.conf and no one is getting locked up for malicious intent.
What if you intentionally or unintentionally took down a server that controlled important infrastructure which people depended greatly on? Flood warning system for example ?
Grow up.
Either way, I still think you're in the wrong, kind of like checking out a naked person getting changed because they accidentally left their blind open. It was available, maybe it was clever, but it's a strange way to learn how to code. Why didn't you just buy a coding book, or borrow some from the library ? Was the code really of good quality if the server was configured so badly?
Obviously we have a difference of opinion and that's ok.
Sure, you can argue that they were not supposed to read the code, so they shouldn't have. But without some tangible harm I don't see why we're supposed to disapprove of it. Maybe allow some hacker spirit while posting on Hacker News :-)