What Copilot means for open source
matthewbutterick.com
matthewbutterick.com
I don't know if the author understands how these transformer models work, but this would a impossible task in Byzantine complexity. The way these models work is by outputting a probably distribution of likely "token" embeddings given an input prompt. The output involves basically inverting a word2vec (probably beefed up with programming language keywords and other bells and plausible search techniques I don't have access to the details of).
This model was of course trained with real code, but you can't attach to the output any meaningful information from the gradient you get from the sample. It's a very messy computation to even think to write down (attaching a percentage that 1 training example affected 1 particular output), much less come up with a simple interpretation of.
You can make a model that is trained only with BSD and MIT code, and IIRC/IIUC the result can be used with any license, including proprietary code.
You can make a second model that is trained only with BSD, MIT and GPL2 code (and perhaps GPL2+), and IIRC/IIUC the result can be only with GPL2 code.
You can make a third model that is trained only with BSD, MIT and GPL2+ and GPL3 code, and IIRC/IIUC the result can be only with GPL3 code.
AGPL, Apache, WTFPL, ... Just add them to the correct model or create a new model for them.
BSD and MIT license still require attribution of the used source code (there exists a MIT No Attribution License, though: https://en.wikipedia.org/w/index.php?title=MIT_License&oldid...).
Which in this scenario could be accomplished by starting that “this software contains algorithmically generated code that was trained using MIT derived code from X, Y, and Z.
This attribution list would be very long, but it could be done.
I think the interesting thing here would be to compare the generated code with the original training data after the fact. If the generated code does not already exist (with some degree of similarity), then it would be okay to use. Otherwise, if it did already exist, you’d then need to attribute the authors of the chunk of code — which you’d now know.
I’m not discounting the hurdles involved, but there might be ways to overcome them.
On the other hand, open source licenses haven’t been tested thoroughly in court, so maybe the result would be to throw out such attribution requirements entirely. The main relevant case to date is https://en.wikipedia.org/wiki/Jacobsen_v._Katzer
In reality through, we'll all just end up using the MIT No Attribution License copilot.
Great, so GPL code would be left alone. I don't see how that's a bad thing.
Surely the solution would be to give credit to every author from the training corpus. I am looking forward to the 10 000 lines of copyrights in every header. :P
If Microsoft had trained it on its own code, there would be no such problems. Surely a company as large as Microsoft has produced enough code over the years to create a large enough training dataset.
I keep seeing this sentiment from the GPL/"laundering" side of the debate.
Believe me, Microsoft wouldn't have released this thing (after what, 6 months of beta testing?) if they thought they had any "problems" at all.
I'm not saying I don't sort of agree with you, but is there no room for what's actually _likely_ to happen in this debate? Because as best as I can tell, they aren't going to see any real legal issues from this.
(There's also an option to remove generations that result in a collision with actual GitHub code, just fyi)
I feel like when the singularity happens HN is going to be flooded with programmers mad that they got automated away despite it very much being one of the primary goals of computer science and software engineering. This stuff is a kind of just a fact of life now.
Salesforce trained models (on GitHub) competitive with copilot without needing to own GitHub. I would spend less time worrying about how to lawyer up and more time figuring out how you're going to adapt to these new tools. That's the gig.
The simply way to test the legal theory behind copilot would be to write a AI that write music notes, using music scraped from youtube or any other large music library. The idea that one can train on "public available material" and produce algorithms that output large chunk of copyrighted material is a bit untested in court, but go against the wrong target and we will quickly see a response. We have actually seen some traces of this with news bots that scrapes news site and produce "novel" interpretation of existing news, especially sports news.
No, we're not. Further, Amazon just announced a similar product and Salesforce has literally _released weights_ for their code models. You can't put the genie back in the bottle.
Actually enforcing any action when the representations are learned rather than hard-coded just seems impossible to me. They have a check box that removes any predictions matching existing code - that basically makes it impossible to discern the source since this will be based on some subjective "semantic closeness" BS.
They would. Did you just forget TAI? Microsoft didn't consider 4chan would train her to be the ultimate racist.
Do you work in marketing? Do you program?
I'm a machine learning engineer, amateur researcher and open source contributor. Before that I was a software engineer for 8 years.
As a person that teaches and conducts research with these models, there is no reason why a tool like Copilot has to be a purely parametric, generative transformer like GPT-3. Where attribution, as you describe it, is nearly impossible.
For example, if the model was to use a retriever component to obtain specific pieces of concrete code from a database (where the licensing for each piece is known) conditioned on the original source code context; then generate its output based on this, the context in your source code, and pre-trained parameters, then it could theoretically at least satisfy Butterick’s request and probably be more akin to how a human programmer operates.
This still does not rule out possible legal issues entirely as the pre-trained parameters are opaque, but it certainly makes it less problematic.
Alternatively, there is active research on attributing a given output to a transformer’s training data. But it is still very early days and frankly it is very “foggy” as to what degree this can be done.
Lastly, I have seen a bunch of comments elsewhere alluding to it being somehow sufficient to reduce the level of verbatim copying. Sadly, it will not, as even if you for example replace the variable names it is still a copyright violation. Just like if you manipulate the RGB space slightly for an image. Determining fair use, etc without a license is something that today can only be done in court, regardless of how we of a heavier technical disposition may feel about it. After all, that is exactly why we have explicit licenses in the first place!
That's what Copilot does with extra steps. I don't see why it should be different just because Microsoft mixed the code snippets in a blender to obscure what they're doing. It's code laundering.
using copilot is like subcontracting to a company who you know has no institutional qualms about their employees liberally copy/pasting from a known corpus of github repos. you know these employees don't always literally copy/paste, but all they do every day is read the corpus, so their code is always going to be basically derivative. you can optionally ask that the company avoids matches to existing code, which you know means that before sending you the code, the boss will fuzzy search their code against the corpus, asking their employees to write it again if a match is found.
so given that this is the way the company works, and that you know this is the way they work, your task is to decide what your ethical and legal liabilities are.
actually i'm sure you do, you just don't think of it this way. The education you received to learn to code would consists of such snippets and examples. Then you sell your skill as a programmer, which consists of you recalling past experience and knowledge, and transform that into the final piece of code.
Copilot is merely doing something like that, but way less sophisticated than a human brain.
I think this will require looking at some actual examples to try to judge, and figure out how to articulate the basis for the judgement.
Hm. That begs the question what good is it if it deals in such small bits? I guess that means if the snippets are large enough to be of any value at all, then they are automatically big enough to be damning.
Analogy: the law has no problem with you memorizing the license plates of cars you see as they pass your street; however, building an automated system for recording and storing license plates in a database is a different story.
Attribution would still be a problem though. How to attribute generated code that's a random mishmash from thousands of inputs? The only solution can be to attribute all of them by doing some sort of reverse match. This "IP scanning" must clearly also be provided by Copilot and not be offloaded to the user.
TBH I would have expected that the "evil geniuses" who came up with Copilot would have thought about such obvious pitfalls beforehand. It's not much of an achievement to innovate by ignoring the law.
Assuming the generated code does not exist and is not part of the training set (these are two giant caveats), how would this be different from me reading GPL code to learn how to do something, and then rewriting it in my own words?
If I have not copied, but learned from GPL code, is my new code GPL? No, it’s not a derived work in terms of copyright. Am I standing on the shoulders of giants? Yes. But so long as that concept is expressed in a novel way (not just changing variable names), then it isn’t a derivative work.
I haven’t worked with copilot at all to know how verbatim the results are, but it is theoretically possible to train a model with GPL code and not get GPL code generated out the other side. (Again, here be dragons).
No different - that is also not allowed. The idea of clean room reverse engineering exists exactly for this reason.
My understanding is the GPL based on copyright - and that copyright only protects a specific expression, not a general idea or concept. To suggest that if something is copyrighted, humans cannot learn from it and generate thier own material seems to be absurd.
But will it be capable of generating a copyright infringement? Let's imagine I want to implement an algorithm to find the shortest path between nodes in a graph. Well I remember from my MSc that you probably need to implement it with Dijkstra's algorithm, and I've implemented this a few times over the years in kata and leetcode. Now let's say I'm programming Rust, which I don't really know the syntax for so I go read some GPL code while browsing the net for syntax.
Do I need to state that my implementation of Dijkstra's algorithm is 10% GPL, because I spent a few minutes reading syntax of a GPL file? What if someone else's copyrighted code looks 90% similar to mine, because hell there's only so many ways to implement it, can I get sued because there is 90% overlap and I might have looked at this code?
These systems aren't even capable yet of generating more than a a few functions, much less a coherent library. Even then, as they generalize more and gain more expressive power they will be less likely to copy in the training data to the output now, a possibility that I consider quite remote even given a relatively primitive Co-Pilot.
When you write your implementation of Dijkstra's algorithm then you hopefully using human intelligence to write it, not an illusion of intelligence. If its original then its 0% of whatever GPL file you happen to have read at a previous date. If you copy something verbatim without using human intelligence, then its 100% GPL.
True artificial intelligence == magic. if not magic then math.
At least according to Google’s monetization policy when it comes to music.
In particular an algorithmic concept cannot be copyrighted but a description of an algorithm in code can be. I think that many programmers believe that the algorithm is the copyrightable bit because that is the “hard part” compared to say the variable names and comments which are the easy part. But copyright cares much more about the latter than loops and conditionals, which have much less creative value. Mathematical expressions cannot be copyrighted.
But of course this wouldn't catch copies where variable names have been changed, etc. One thing that I think would be really interesting (but hideously computationally expensive) is to compute an embedding of each chunk of the training data and then query for the k nearest neighbors when Copilot generates some code, so you can see what the closest snippets in the training data are and evaluate for yourself if they're too similar.
Checking generated code for similarity against training set is possible, and is now done by both Copilot and CodeWhisperer. But it'll include code that just happens to be similar, even if that code had no influence on what the model generated.
I was finding the copilot trial was pretty good at reading something like an adjacent CSV file and building out a struct with proper data types from the data in it but beyond writing CSV->DB migration stuff I didn't really use it for a lot
Sidestepping the whole license violation issue, using GitHub copilot is the equivalent of using an IME to type Chinese - I may not remember exactly how to write the character but I'll quickly recognize it when I see it. It's the difference between being able to write traditional Chinese versus the ability to read the characters.
Anytime that I don't have to context switch and alt tab to look something up on MSDN or in Stack overflow is an absolute freakin win for me as a developer.
It kills me that people can't seem to comprehend this.
because it's not a good comparison. Code is not language because language need not be syntactically correct, and people generally can infer correct semantics even from badly mangled communication.
These automated coding solutions generate syntactical and semantic errors that another machine will not understand, and even more importantly, it generates kinds of errors that people are not accustomed to.
Copilot is more like an automated car that goes off into random directions in ways that even human drivers would not, and you never know when. And when humans have to interact with black boxes whose behavior they cannot anticipate, small errors are not merely small errors, they create significant tension and uncertainty that requires almost constant attention.
That’s more just a product usefulness complaint rather than any kind of ethics. If it’s working for others, that’s good.
One of the hardest things to come to terms with as a budding young software dev is that most real world code is as wide as an ocean and as deep as a puddle, and the occasions where it's harmless (let alone actually helpful) to be clever are few and far between.
It works even better in languages like Haskell or Ocaml, where often there is only a couple ways (often only one way) valid code could be written once you type part of it out and if it does spit out invalid code you get an instant compiler warning.
1) As automatic documentation in a well-typed application. Autocomplete shows you the public properties or methods available based on what data you have started typing, and it's not guessed but guaranteed to be valid assuming your types are correct.
2) As dumb autocomplete, where it just tries to guess based on the symbols in the document, and save you some keystrokes/remembering.
3) As the Copilot style autocomplete where it will finish the entire line, and not just the next token—with the downside that it's really just guessed.
I can see where 2 and 3 can be avoided, but 1) is really invaluable and often almost required in some strictly typed languages. I remember the days of writing PHP where I had to google for the function signatures every time because of inconsistent naming, and lack of typing. Memorizing these things is honestly a waste of a programmer's brain space. But now that I write exclusively in typed languages, auto-complete in the form of 1) has been invaluable not for the purposes of remembering names of things but being able to see what functions/methods are valid and can be used. No need to look up documentation manually anymore.
Automatic completion or suggestion popups are super annoying, though. Even worse, some IDEs will dynamically reformat the code including the line I'm still typing. Feels like Clippy in MS Word but instead of showing options it automatically just does something random without asking. Efficient keyboard shortcuts make unreliable "helpers" redundant.
Whereas the difference between two developers one who remembers the exact parameters for some random arbitrary library and the other who relies on autocomplete is probably negligible. There's a huge difference between rote memorization and fluid intelligence.
What's more important is your ability to solve problems and develop new algorithms and you can do that in pseudo code without memorizing a single pointless parameter list.
If I read the code for, I don't know, some GPL-3 library and then write my own MIT-licensed version--that's totally fine.
A programmer can read strictly licensed code and then use that knowledge to write their own non-strictly licensed code.
Copilot is not different from a human. It has knowledge & it uses it. It isn't copy-and-pasting (there are 1/2 edge cases where it is; but for the most part it's new ideas).
It's the same thing as saying "Dall-E 2 is plagiarizing art".
When I write a quicksort algorithm, I don't give any attribute to the code I saw for the algorithm in some random library.
Fundamentally, there's no real difference between Copilot and a human. I've watched Copilot write crazy lodash one-liners that were clearly contextual to my code.
I think what is fundamentally happening here is that older people / people of the last generation are realizing that just as sys admin jobs / etc. are going away, soon many rote coding jobs will be taken away since Copilot will automate them. And that's fine, but it is producing backlash which comes in the form of licensing issues.
Basically, there's no difference between Copilot and Dall-E and it's pretty clear that Dall-E has no licensing issues, thus Copilot should also be in the clear.
This is incorrect. Humans are capable of creating. Copilot is merely capable of regurgitating.
> If I read the code for, I don't know, some GPL-3 library and then write my own MIT-licensed version--that's totally fine.
It's actually not totally fine and there have been many many many court cases over this sort of thing, both with non-commercial and commercial licenses. The whole concept of clean-room implementation exists as a defense to this.
> When I write a quicksort algorithm, I don't give any attribute to the code I saw for the algorithm in some random library.
If it's substantially similar to the library's, it's entirely possible you're violating their license terms and/or copyright.
> Basically, there's no difference between Copilot and Dall-E and it's pretty clear that Dall-E has no licensing issues, thus Copilot should also be in the clear.
I'm not sure this is correct. If DALL-E began outputting verbatim copies of other people's works, they could very well be sued over it. Similarly, if it produced trademarked symbols like the Nike swoosh, the Golden Arches, or the Starbucks logo, it's not like those aren't going to get you sued.
Infringement is about the produced thing (code block or image or whatever else) and its use, not the method of generating it.
I claim copilot is capable of creating. When copilot writes an amazing lodash one liner in my code --that isn't regurgitating an existing lodash snippet, it's creating something new to fit my use case. This is undeniable and there is nothing to argue here. Copilot regularly looks at my code and writes new code that is highly specific to my existing code. It's honestly better than me at languages I don't know well (like C).
And yes, 0.1% of the time copilot is spitting out existing code verbatim; the other 99.9% of the time copilot is not overfitting and is synthesizing new code. Luckily most companies don't needs clean room implementations.
Copilot writes code that is adapted to my existing code base. This is quite obviously and undeniably not just regurgitation because my codebase is unique to me.
Lastly copilot is a better programmer (locally) than many of my peers. It can write better lodash one liners, amongst other things, and while that's embarrassing it's true.
I posted elsewhere in this thread talking about my experiences (both old and recent) using it. Dismissing a reply you don't like simply because of a (terrible, given the tool is very available) assumption is just a poor quality response.
> I claim copilot is capable of creating. When copilot writes an amazing lodash one liner in my code --that isn't regurgitating an existing lodash snippet, it's creating something new to fit my use case. This is undeniable and there is nothing to argue here. Copilot regularly looks at my code and writes new code that is highly specific to my existing code. It's honestly better than me at languages I don't know well (like C).
> And yes, 0.1% of the time copilot is spitting out existing code verbatim; the other 99.9% of the time copilot is not overfitting and is synthesizing new code. Luckily most companies don't needs clean room implementations.
Where do you get this figure of 0.1%? I'm not aware of anyone having studied it, and absent that the figure seems entirely fabricated and could be higher or lower. Indeed, the way it works via prompting suggests that an overall percentage is irrelevant if 100% of the time you ask it for specific things it generates copyrighted or strictly licensed code. If you have references though, I'm interested.
> Copilot writes code that is adapted to my existing code base. This is quite obviously and undeniably not just regurgitation because my codebase is unique to me.
Using your code to show you code you might likely write isn't "creative" and is exactly regurgitating what it's seen. Is a Markov chain "Creative"? That's essentially what you're describing here.
> Lastly copilot is a better programmer (locally) than many of my peers. It can write better lodash one liners, amongst other things, and while that's embarrassing it's true.
You've used lodash one-liners as an example a couple of times now. Why? Why is that a litmus test for a good programmer? What makes the copilot generated ones superior? Do you have examples?
My experiences with copilot as I've shared elsewhere are that it produces about 90% of the time code that needs to be debugged, doesn't quite fit coding conventions we use, and often doesn't do exactly what I'm looking for. It's fine if you're using it to stub in specific kinds of boilerplate or using it to generate a function to do some very standard math thing that exists or should exist in a library somewhere.
I talk about Lodash one-liners because it makes it pretty obvious that Copilot is not just copy-pasting code (which would be copyright infringement). It's quite unlikely (even if we consider all variables to have the same name), that Copilot is copy-pasting an exact copy of some other snippet, given that the snippet is very specific to my problems. (I'm not asking it to write a generic math function, I'm asking it to use lodash on data structures in my codebase to accomplish a very specific outcome). By quite unlikely I mean (1 / (2 ^ 32)), if we consider the one-liner to be composed of 32 different AST nodes.
> Is a Markov chain "Creative"? That's essentially what you're describing here.
Have you read the paper that (eventually) inspired Copilot? "Attention is All You Need". It's not a Markov chain. There's many, many, many different steps / layers. I think if you know & understand the fundamental building block that is responsible for it working (the "transformer"), then a lot of the worries around plagiarism go away.
Also. What is coding if not a search problem over a very large space? I'm searching for the next N lines to write over the space of all possible lines. To guide my search I use things like my prior knowledge.
In my head, this knowledge is encoded using neurons. In Copilot, the knowledge is encoded in parameters. Abstract concepts in my head are encoded in neurons. In Copilot's "head", abstract concepts are encoded using "embedding vectors" of 512 bytes (maybe more / less, not sure). Yeah, maybe I use more than 512 bytes to encode a concept, but still, I don't see a huge difference. If I'm not plagiarizing, then I can't imagine Copilot is (except for in the 0.1% of cases where it's over-fitting)
I suggest you read the Sony v. Connectix appeal veredict.
Is the risk that you'll be sued because some developer unintentionally used fully-reproduced code worth using copilot?
There's a reason people sometimes clean room document some code, then have a different set of people who never read the source re-implement it with out ever having seen the source. To avoid these kinds of issues.
I don't think that's always fine.
> A programmer can read strictly licensed code and then use that knowledge to write their own non-strictly licensed code.
If you've ever read the Windows source code, you're never allowed to contribute any code of your own to Wine or ReactOS.
Copyright should apply to large and whole pieces of work only. The whole of a painting should be copyrighted. The style and technique should not. Same for code. Windows as a whole should have copyright. The snipit that handles a mouse click should not.
Fundamentally, it's absolutely OK to allow humans to do the thing X and to prohibit AI from doing it. This is what I hope will happen here (though it probably won't).
>soon many rote coding jobs will be taken away since Copilot will automate them. And that's fine
Good luck finding enough non-rote jobs to re-employ those developers. It will be an interesting reflection though when "just learn to code" turns into "just learn to lay bricks for $10 an hour".
When it was books, films, videos, news articles, journal articles, blog posts, photographs, extremely private personal data, music, art and other creative works being hoovered into the giant AI exploitation cloud in the sky, and non programmer creatives complained, they were generally labeled as luddites, trolls, whiners, maladapted (“learn to code”), and just generally ignorant worthless trash.
Now that it’s happening to you oh boy do you care.
Why care about variable naming since AI will generate most names anyway, why care about common libraries if you can always get good implementation for specific use case in a second or two?
I know these small details don't matter in the big picture, code exists to do something and if it does, it doesn't matter how it was made.
Copilot introducing "learned" security issues, for example, is a problem separate from the copyright risk.
Currently, it's often a matter of updating a dependency to patch things away. How do you do that when you barely understand the code it helpfully generated for you?
People will jump to say "well that's using it wrong" but the reality is if you allow this in your org people will use it this way.
With verbatim code copies out of the way, I don't see any basis to consider code produced by Copilot copyrighted. So unless you have some problem with the fact that portions of your code will be public domain, I don't see any reason not to use it.
And one more thought. The author gives an example of using Copilot to list prime numbers. That's not a good use for it. Copilot and similar systems are primarily useful for writing boring boilerplate code, saving your time for more involved parts.
That is literally the first kind of usage promoted on the main landing page for Copilot. It shows Copilot filling in the body of function definitions in three different languages.
Assuming a use case, why in earth would I trust such a Trojan Horse from Microsoft, seeing as how it’s likely serving its master in intended ways I can only guess at?
It’s not a useful tool, imo, but then I don’t use IDEs or autocomplete.
Maybe I’m not the target developer.
It’s a duck problem. Quacks like trouble, smells like trouble, looks like trouble, comes from an arch-enemy of FOSS that bought a major FOSS hub and is now seeking to do something with its purchase, which I’ll wager isn’t a good upright wholesome thing.
Why invent things when I was perfectly clear?
You also write "Quacks like trouble, smells like trouble, looks like trouble" which is simply baseless. This particular passage triggered my observation that you simply hate it without any justification.
I totally appreciate that you might not find Copilot useful, but your comment went farther than that.
If you have a looser definition, I don't see why you wouldn't be able to copyright code that you generated by using Unreal Engine's blueprint to C++ utility, or any other tool that assists in transpiling code. 99% of the web consists of transpiled JS these days.
> : to define or originate (something, such as a mathematical or linguistic set or structure) by the application of one or more rules or operations[0]
Creating something from the application of rules or operations sounds just like what a compiler does. Whereas to process something is:
> : treated or made by a special process especially when involving synthesis or artificial modification
So if anything it sounds like you're asking if processed code is copyrightable. But this is just quibbling over pedantry. My original point stands. AI is just following a predefined set of rules to transform your request into code. It's a program, just like a compiler is a program. So it would be really hard to say that AI generated code is non copyrightable, but compiler/transpiler/fuzzy generated code is copyrightable.
I feel like this understates the wild west nature of software in non-tech Fortune 500
Copilot never generated anything for me that would justify any kind of copyright.
It'll be highly dependent upon the type of output it generates.
I've played with it again recently because of the updates around disabling open code and such, and it still strikes me that it's not worth the potential risk.
Even setting aside the "could I get sued for this" question, the value just isn't there; half the time functions are just poorly written, inefficient, or buggy. For very simple actual math stuff (e.g.: a function to replicate any moderately complex spreadsheet function) it seems to work well. It seems to really struggle with common boilerplate stuff like simple HttpServer in Java, Flask blueprints, and so on. It may be bias due to the projects I've worked on, but I don't do a lot of my own implementations of calculating compounding interest rates, incredibly simple array slicing, or testing if a given number is prime.
They may find that SCO owns Linux for all we know… vote. Yesterday would have been ideal, but moving forward, remember to vote and remember which party hates common sense
I think the HTTPS proxies necessary to reach the outside would block the communications necessary for Copilot to work.
Maybe all AI secretly wants to kill us.
I saw the post some minutes ago and it was with posted with the original titles. I guess the mods changed it. The part about wanting to kill the writer is an exaggeration. It's usual to changed it to the subtitle or a relevant sentence, but without too much cherry picking (the last part is not very clear).
Someone made a tracker to for these changes https://hackernewstitles.netlify.app/ HN discussion https://news.ycombinator.com/item?id=21617016 (366 points | Nov 23, 2019 | 94 comments) Most of the changes make sense (if you forget the infamous case of the asteroid/rocket part).
A man can dream.
On the practical level I agree with the part advising caution to those that might end up embedding an identifiably licensed snippet in their codebase via copilot. I also agree that copilot users plagiarizing significant chunks of GPL code for profit is immoral. This needs to be prevented.
I also share the frustration stemming from big companies leveraging their disproportional access to data and resources for profit given that the greatest value of these models is precisely the open source code it is trained on.
Ultimately though, what I care about is the potential for building better tools. LLMs potentially offer paths towards genuinely new forms of human-machine interaction, and I don’t want that exploration to be suffocated by legalism.
Wealthy corporations are never going to be "suffocated by legalism" — they can afford to do their research privately. (And many still do.) The issue here is that Copilot is being foisted into the agora seemingly without sufficient (or maybe any) scrutiny of its legal consequences.
More broadly we are seeing a norm emering where there is so much hype chasing AI that these wealthy corporations (see also: Tesla, of course) have a huge incentive to push their experiments into the public sphere prematurely, simply to assert their primacy.
BTW this technique of front-running regulatory scrutiny can still backfire. If the initial public experience with an emerging technology is sufficiently bad, it can poison the acceptance forever. IOW, you can be suffocated by your own hubris faster than any external legalism.
What worries me the most is the effect the public backlash towards these big companies can have on smaller actors that could enter this space in the near future. In the past we’ve seen open source projects like GPT-J come together to fund and reproduce closed models, and if we’re not careful to be nuanced in our criticism of big-tech frontrunners we might end up poisoning the waters enough to deter small actors without a dedicated legal team.
Copyright law is ultimately designed around humans as the only kind of actor. In an ideal world we would sit down and think about the way non-human learners should fit into this system and the balance of tradeoffs we want those laws to aim for. I hope that happens someday, but until then I hope we can cultivate a world where small actors are able to experiment with these technologies without fear of legal action.
That’s why it bothers me to see people arguing that language models should be thought of like human programmers making derivative works, even suggesting that we should require attribution for all generated outputs (i.e. the entire training set, always). That helps nobody, except of course big companies with infinite manpower.
Should someone ask a senator's chief-of-staff to send a strongly worded letter (on behalf of the senator) to ask Microsoft to publicly answer these pressing questions? We're several months in and the silence is deafening
Call me a pessimist but it seems really unlikely that AI will cause problems so much as continually increasing automation. It makes such terrible choices around qualitative decisions and increasing existing AI to a general solution pretty much always fails.
So yeah, most complains are quite right, how you enforce them? You could go against, maybe 10-50, but you'd still have the rest of the planet full of non-complainers (the more popular, good, efficient you code was at the moment Copilot learnt it).
So the problem isn't what's right or wrong, but it is that short of shutting down Copilot, most of the alternative solutions have no enough impact, not even close to the change Copilot is creating right now,
How singularity begins? It could have already started.
And that a lot of github's autofills will read "please see my sourcehut repository for details"
If it has issues in the current framework of Copyright law, then that is yet another reason to change that framework.