Such projects should always be opt-in, not just because it is the law but also because it is common sense and the right thing to do from an ethical perspective.
Such projects should always be opt-in, not just because it is the law but also because it is common sense and the right thing to do from an ethical perspective.
Fyi... Google Books (scanned and OCR'd books) eventually won against the authors filing lawsuits of copyright infringement. So there is some precedent that courts do look at the "utility" or "sufficiently transformative" aspect when weighing copyright infringement.
https://www.google.com/search?q=google+books+%22is+transform...
But courts in Europe may judge things differently.
The thing that surprised me about that ruling is that it was deemed final without a chance of an appeal.
An libraries are very special entities.
I still think "training is fair use" still has a leg to stand on, though. But it doesn't save GitHub Copilot because they're not merely training a model; they're selling access to its outputs and telling people they have "full commercial rights" to its outputs (i.e. sublicensing). Fair use is not transitive; if I make 100 Google Books searches to get all the pages out of a book, I don't suddenly own the book. There is no "copyright laundry" here.
If that's the case, we need to serious re-consider how we reward Open Source as a society (I think that would be fantastic anyway!) -- we have people producing knowledge and others profiting directly from this material, producing new content and new code that's incompatible with the original license.
You make GPL code, a make an AI that learns from GPL code, shouldn't its output be GPL licensed as well?
I think meanwhile the most reasonable solution is that an AI should always produce content compatible with the training material licenses. So if you want to use GPL training sets, you can only use that to create GPL-compatible code. If you use public domain (or e.g. 0BSD?) training sets, you can produce any code I guess.
No. It's called copyright, it is enforceable, and that's the control.
GPL source code is available everywhere, in all formats, in textbooks, on CDs, on websites, but it is still gpl.
And Microsoft doesn't get to scrub the license.
GitHub's argument* is not that they're following the license but that the license does not apply to their use. So they would continue to ignore any provision that says they can't use the material for training.
Previously discussed: https://news.ycombinator.com/item?id=27740001
Moving off GitHub is a better step at a practical level. But again they claim the license doesn't matter, so even if it's hosted publicly elsewhere they would (presumably) maintain that they can still scoop it up. It just becomes more work, for them, to do so.
*Which is completely wrong in my opinion, for the record
I think, for the desired outcome to occur, you should instead ask:
You write close sourced code, then a make an AI that learns from that code, shouldn't its output be licensed as well?
Ask the above, and suddenly Microsoft will agree.
> Ask the above, and suddenly Microsoft will agree.
Does Microsoft actually agree? Many people have posted leaked/stolen Microsoft code (such as Windows, MS-DOS 6) to GitHub. Microsoft doesn't seem to make a very serious effort to stop it – sometimes they DMCA repos hosting it, but others have stayed up for ages. They could easily build some system to automatically detect and takedown leaks of their own code, but they haven't. Given this reality, if they trained GitHub Copilot on all public GitHub repos, it seems likely that its training included leaked Microsoft source code. If true, that means Microsoft doesn't actually have a problem with people using the outputs of an AI trained on their own closed source code.
It's a surprisingly subtle distinction.
EDIT - if I squint hard enough in exactly the right way, there's a sense in which CoPilot etc aligns perfectly with the goals of the free software movement. A world in which you can use it as a code copyright laundry might be a world where code is actually free.
Is that any weirder than bizarre legal contortions such as the Google/Oracle "9 lines of code"? Or the whole dance around reverse engineering: "It's OK if you never saw the actual code but you're allowed to read comprehensive notes from someone who did"..?
There's a ton of examples like this. Tell me with a straight face that there's a clear moral line in either copyright or patent law as it relates to software.
IP is a mess and it's not clear who benefits. Is a world where code isn't subject to copyright so bad?
It is specifically the idea of using copyright to eat itself that is harmed by AI training. In the world where we currently live in, only source code can be trained on. If I want to train an AI on, say, the NT kernel; I have to decompile it first, and even then it's not going to be good training data because there's no comments or variable names to guide the AI. The whole point of the GPL was to force other companies to not lock down programs and withhold source code, after all.
Keep in mind too that AI is basically proprietary software's final form. Not even the creator of an AI program has anything that resembles "source code"; and a good chunk of AI safety research boils down to "here's a program you can't comprehend except through gradient descent, how do we design it to have an incentive to not do bad things".
If you like copyright licensing and just view the GPL as an exception sales vehicle, then AI is less of a threat, because it's just another thing to sell licenses for.
If the output (not just the model) can be determined to be a derivative work of the input, or the model is overfit and regurgitating training set data, then yes. It should. And a court would make the same demands, because fair use is intransitive - you cannot reach through a fair use to make an unfair use. So each model invocation creates a new question of "did I just copy GPL code or not".
Overall it feels like it's a bit too much of specialized learning on GPL/Copyleft code to be fair. It's not like a human that reads some source code and gets an idea how it works. It's really learning code from scratch on Copyleft code, without which it would likely perform much worse and not generate a number of examples. It's not just copy-paste, but it's closer on the spectrum to copy paste than just super-abstract inspiration to feel fair.
As others have said, I don't think it would be fine (specially from big companies pov.) to decompile proprietary code (or just grab publicly available but illegal to reproduce code) and have AIs learn from that in a way that seems different in scope and ability to human research and reverse engineering.
I think we need a good tradeoff that isn't ludditism (that would reject a benefit for us all, i.e. that is good for everyone), but that still promotes and maintains open source software. In this case it's really a public "good" that's being seized and commercialized, that doesn't seem quite right: make copilot public, or use only permitted code (or share your revenue with developers -- although that would seem more complicated and up to each copyright holder to re-license for this usage). I remember not long ago MS declaring Open Source was a kind of "Cancer", now they're relying on it to sell their programming AIs. I personally think Open Source is quite the opposite of cancer, it is usually an unmitigated social good.
Much of the same could be said for the case of artists an generative AI art.
And this isn't even starting on how we move forward as a society that has highly automated most jobs and needs to distribute the resources and wealth in a good way to enable greatest wellbeing for all beings.
Is that new? If I include some excerpt from copyrighted material in my own work and it's deemed to be fair use, that doesn't limit my right to profit from the work, sell the copyright to someone else, and so on, does it?
But if you read the source code of 100 different projects to learn how they worked and then someone hired you to write a program that uses this knowledge, that should be legit. I'm not sure if the law currently makes a distinction between learning vs. remixing, and if Copilot would qualify as learning.
But, for the general case, the argument still stands. I have looked at GPL code before. I might have even learned something from it. Is my brain infected? Am I required by law to license everything I ever make as GPL for the remainder of my days?
I mean in essence Github is a library; they did have a license to a point to do with the code as they pleased, but they then started to create a derivative work in the form of an AI, without correctly crediting the source materials.
I mean I think they made a gamble on it; as far as I'm aware, AI training sets were yet unchallenged in a court of law, so legally not fully defined yet. These lawsuits - and the ones (if any) aimed at the image generators, using CC artwork from e.g. artstation - will lay the legal groundwork for future AI / ML development.
Entities like the Internet Archive skate by (at least before their book lending stunt during COVID) by being non-profit and bending over backwards to respect even retrospective robots.txt instructions, meaning that it's not really worth suing them given they'll mostly do what you ask anyway.
But I guarantee you that if I set up a best comic strips of all time library I'll probably be in court.
I find this to be a very appropriate analogy. If Google had done such a thing, they would be facing the same kinds of lawsuits that Microsoft is facing now. And despite Microsoft's money, I don't see how they can wiggle their way out of this one. They basically ignored the license terms and attribution requirements of the authors. Something Microsoft would never stand for, if "the shoe was on the other foot".
They did appeal it. SCOTUS declined to hear the case.
https://www.nytimes.com/2016/04/19/technology/google-books-c...
Google Books retains the bibliographical information so you can properly cite the authors or contact them for permission to use their material.
And Google Books does not automatically write new books for you that you can then send off to Penguin Books or self-publish on Amazon.
That sounds to me somewhat close to "if I take an FFT of each of those copyrighted images, glue them together, and sell this as a picture, is that a derivative work?" - I'd say yes, or perhaps even a different encoding of the original work, since you can reverse the frequency domain representation and get the original spatial representation - the original images - back.
Sometimes parts of the original works are still encoded, which we've seen when some code is reproduced verbatim, and I'm sure that happens to people as well, ie. they see some algorithm and down the road have to write something similar and end up reproducing the exact same thing.
Once they iron out those wrinkles, it's not clear to me that a large language model is a directly reversible function of the original works. At least, not any more than a human learning from reading a bunch of code and then going on to have a career selling his skills at writing code.
Edit: by which I mean, LLMs are lossy encodings, not lossless encodings.
They're not, but the "giant table of token frequencies and associative keywords" reminded me of doing FFT on images, and I wanted to communicate the idea that transformations like this can actually retain the original information, and reproduce it back through inverse transform.
> by which I mean, LLMs are lossy encodings, not lossless encodings
Exactly. And while I doubt most training data is recoverable, "lossy encoding" is still a spectrum. As you move away from lossless, it's not obvious when, or if at all, the result is clear from copyright of original inputs' author. Compare e.g. with JPEG, which employs a less sophisticated lossy encoding - no matter how hard you compress a source image, the result would still likely retain the copyright of the source image author, as provenance matters.
(IANAL, though.)
I'll just finally note that LLMs are not lossy encodings in the same sense as JPEG. LLMs are closer to human-like learning, where learning from data enables us to create entirely new expressions of the same concepts contained in that data, rather than acting as pure functions of the source data. That's why this will be interesting to see play out in the courts.
You could say a higher order program can "just" be transformed into a first-order program via defunctionalization, but I think the expressive difference is in and of itself meaningful. I hope the courts can tease that out in the end, and we'll see if LLMs cross that line, or if we need something even more general to qualify.
Interesting analogy, and I think there are a couple different "levels" of looking at it. E.g. fundamentally, they're the same thing under Turing equivalence, and in practice one can be transformed into the other - but then, I agree there is a meaningful difference for humans having to read or think in those languages. Additionally, if those are typical programming languages, you can't really have the code in the "weaker" language self-upgrade to the point the upgraded language has the same expressive power as the "stronger" one. If the "weaker" one is Lisp though, you can lift it like this.
In this sense I see traditional compression algorithms - like the ones we use for archiving, images and sound - to be like those typical weaker languages. There's a fixed set of features they exploit in their compression. But human learning vs. neural network models (or sophisticated enough non-DNN ML) is to me like Lisp vs. that stronger programming language, or even Lisp vs. a better Lisp - both can arbitrarily raise their conceptual levels as needed. But it's still fundamentally compression / programming Turing machines.
And if such algorithm is copyrighted, that would be infringing! It doesn't matter if you copy on purpose or by chance.
I'd have thought that's exactly what you can't do with CoPilot.
If you overlap a hundred different FFTs, then the result is likely fine copyright-wise.
These networks are not [supposed to] contain much of the original data. Like the trivia point that Stable Diffusion has less than two bytes per source image, on average.
Stitch them side by side. Yes, this is not how those DNNs work, but the example was more about highlighting that "a giant table of token frequencies" by itself is probably reversible back to original data, or at least something resembling it.
> Stable Diffusion has less than two bytes per source image, on average.
I'm not convinced by this trivia point, though. Stable Diffusion is, effectively, a lossy compression of the training data. Nothing says lossy compression algorithms can't exploit some higher-level conceptual structures in the inputs[0], and applying lossy compression to some work doesn't automatically erase the copyrights of the original input's author.
--
[0] - SD isn't compressing arbitrary byte sequences, it's compressing images - which is a small subset of all possible byte sequences as large as the largest image used in training. "Less than two bytes per source image, on average" doesn't sound to me like something implausible for a lossy compressor that is focused on such small subset of possible inputs, and gets to exploit high-level patterns in such data.
That depends entirely on how many frequencies you're keeping.
> high-level patterns in such data
High level patterns across thousands of images are generally not copyrightable.
I might even describe the purpose of stable diffusion as extracting just the patterns and zero specifics.
Two bytes would only let you uniquely identify ~65k images though, which to me doesn't sound plausible for a lossy compressor.
Curiously, from the article, copyright infringement is not alleged:
> As a final note, the complaint alleges a violation under the Digital Millennium Copyright Act for removal of copyright notices, attribution, and license terms, but conspicuously does not allege copyright infringement.
Perhaps the plaintiffs are trying to avoid exactly this prior law?
Create a for-profit copyright registry for code snippets that are long enough to qualify for copyright protection. You can be the canonical owner of the copyright for a given piece of code! For a premium fee, we can generate and submit a patent on your behalf as well.
Once I have a large corpus (perhaps millions of entries of code, most one or two lines long), I can automatically scan new respositories and send cease and desist letters for violating my client's copyright. Even if a piece of code is very common, that doesn't mean its unoriginal, it just means that there are many people violating its copyright after all. According to the logic of the folks in this thread at least.
Just yesterday, I bought a nice domain name for an idea that's very close to what you mention, monetize on snippets of code.
If you want to team up, hit me up!
This may be one of the reasons why the lawsuit isn’t based on copyright.
But your idea is pretty much what almost all manufacturers do, and have been doing for decades.
In copyright law there is no such concept as "code snippets that are long enough to qualify for copyright protection" or "canonical owners". Quite explicitly, copyright does not give a monopoly over an idea, but merely protects against the unlawful reproduction of an original work.
If you take some snippet from a work in which you own copyright and find that in the world multiple people have somehow managed to write the exact snippet, but they did it independently without copying it from you, then copyright law effectively states the following things:
1) They definitely aren't violating your copyright, and you have no claim on them whatsoever - independent creation is a complete defense to copyright infringement;
2) Perhaps this snippet might be judged uncopyrightable, as the existence of multiple independent recreations is some evidence that it lacks originality and thus would not qualify for copyright protection at all.
There does exist the concept of 'originality' in copyright law, which I was erroneously conflating with length.
https://jacquesmattheij.com/what-is-wrong-with-microsoft-buy...
What do you use instead?
The top alternatives in my opinion are:
- SourceHut https://sr.ht/
- Codeberg https://codeberg.org/
- Self-hosted using Forgejo https://forgejo.org/ (fork of Gitea)
I was self-hosting my code with Gitea for a while but currently I’m using GitHub. Planning on setting up a Forgejo instance in the coming weeks.
At work we use GitLab, but personally it is one of my least favourite platforms, so I am excluding GitLab from the list above.
Forgejo is a fork of Gitea, not Gitlab.
The most likely of which – if this lawsuit ends up winning – is that corporations will have new ways to sue everyone and that the world will be a worse place.
Copyright expansion has never benefited the "little guy" such as Open Source authors, only large entities with deep pockets who can litigate to no end.
People will quote that John Carmack Doom example where it copies the function verbatim, but as far as I can tell that's a rare thing, and it's a function that's been widely copied around without proper licensing; a human could also get it wrong by copying it from github.com/random-person/mit-project with the wrong license (and since then there's also been work to prevent this kind of thing).
Co-pilot isn't unique, or the first AI/ML project to use copyrighted works; all the GPT models use copyrighted works as their input. Some doubts have been raised over the legality of that too, but it's received nowhere near the amount of criticism that Co-Pilot has, certainly not on HN, and I've never seen anyone doubt the morality of it – only the legality.
If you were to go through my public open source code I'm sure you can find stuff that's very similar to some code from my previous employers or other open source projects. Not because I copy/pasted anything, but because my brain was trained on that dataset: you see or write something that works, you face a similar problem a few years later, you write a similar solution.
"Using existing works as input" is common throughout creative works. As Phil Anselmo once said: "with Pantera we took our five favourite bands and ripped 'em off to hell".
People are already getting sued because "that one melody sounds a bit similar to this other melody"; fair use is already widely ignored/disrespected. Much will depend on the exact details, but any win in this lawsuit has a very real chance of empowering that sort of nonsense.
A union of creative minds seems long overdue. It can provide a copyright trust, addressing your concerns, as well as removing the excuse of inability to get permission from n thousand creators used. It can also address matter beyond OSS, such as overreaching employment agreement clauses that assert ownership of everything in your head.
[Possibly 'trust' is more suitable than 'union'. Something like Creative Commons Trust.]
But what prevents Microsoft from harvesting open-source code from any hosting site? What have you gained?
All of those build their models based on the same sources and it's therefore a much more general issue than just one particular company being sued.
But that doesn't make it right in this case and, conveniently, someone has decided to bring suit. The funny thing is that Microsoft depends on Copyright law for their existence and now they want to change the rules to favor them when it suits them. In fact one of the first things that Bill Gates ever did that I remember is bitch about people copying the software that he wrote.
Parking tickets should start at 50% of your yearly take home income. Didn't feed the meter an extra quarter? $10k minimum sounds fair.
Pay is not profits, it's revenue.
I mean, it's obvious that uploading code requires you license the hosting provider a license to host it (which is not singing over copyright); although feel free to argue that the license doesn't or shouldn't extend to CoPilot usage.
So in a way this would void parts of many TOS agreements where you do relicense your User-Generated Content. If we're uploading memes to Facebook, they're gonna have to work out license terms with the copyright holders, not the uploaders.
Suppose person A comitted a crime, that does not mean you are now allowed to profit from someone else's crime
Same will be for GitHub: if people really didn't have the legal authority to bind someone else's code to GitHub's TOS, then GitHub can go after the $x million of users that have uploaded code they shouldn't have.
Ok, but can they go after them in an efficient manner that doesn't end up costing more than it's worth?
Thats doesnt mean. Getty can keep the money
It is a similar case when a single user uploads a movie or game to a pirate torrent site. The site can have a terms-of-use that gives a license to the hosting provider, but naturally the users who upload the content might not have the permission to grant anything to the hosting provider. Depending on how much the hosting provider is or should be aware, hosting the content can still be illegal.
Then, chances are, it's technically illegal to upload those other contributors' code, although if that code is contributed via GitHub itself then the code in the pull request has already been licensed to GH.
It boils down to copyright/DMCA not requiring that hosting providers ensure the code people say they have the rights to is valid at submission, so GitHub now has tons of examples where people themselves lied about the permission when they uploaded code that wasn't theirs, and this will probably be a valid legal defense, at least only for the argument of "does GitHub have the right to use the source in their ML model" (it might really boil down to "are GH's terms vague enough to where nobody thought they the license included the ability to train artificial intelligence").
In the US what it will be is good evidence to support a claim by GitHub that they were an "innocent infringer"--someone who did not know they were infringing and had no reason to believe that they were.
What that does is in the case where the plaintiff seeks statutory damages (which they almost certainly will¹) is lower the lower limit. Statutory damages are normally $750-30000 (amount determined by the court). If a defendant proves they are an innocent infringer that lower limit drops to $200. If the plaintiff can prove that the infringement was "willful" the upper limit goes up to $150000.
Statutory damages are per work infringed, not per infringement, so we aren't talking $200 or so multiplied by the number of copies GitHub distributed. We are talking of a likely award of $200 or so total (plus maybe attorney fees).
¹It is usually way too hard to determine actual monetary damages in cases like this, and actual damages are likely to be quite low anyway, so plaintiffs almost certainly will go for statutory damages.
Can this be said by microsoft? They explicitly chose to not include hidden repositories by their paid customers, likely because they knew that those customers would sue them if proprietary code was used as training data.
Apple seemed to have chosen not to include GPL in the app store for very similar reasons. Their term of service require a permission which is incompatible with the terms of GPL, and knowing that GPL software tend to include multiple rights owners, Apple chose to go the route of not allowing GPL.
And last, authors has requested to have their works removed from the training data. It is part of the lawsuit. Can Microsoft then still claim that they did not know they were infringing?
I believe GitHub would likely be seen as an innocent infringer in that case.
I doubt Microsoft would make that argument. It is more likely they will argue fair use, but by not using closed repositories owned by paying customers, it seems to show that they themselves have doubt about the legal status of using other peoples copyrighted work for copilot.
Or they're worried about leaking secrets, which is a different matter entirely. The amount of copying needed to leak secrets is far lower than the amount needed to commit copyright infringement.
If Copilot is trained on Microsoft's code and accidentally regurgitates a comment, "// for 2024 Xbox", it has done one but not the other.
Copyright infringement doesn't have a fixed size. It depend on context and what kind of information is copied. It demonstrate that copilot has not actually learned how to code (as many people like to claim), but is simply a algorithm for copying code. If it had learned to code like a human it wouldn't divulge secrets.
> 4. License Grant to Us We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
https://docs.github.com/en/site-policy/github-terms/github-t...
I suspect that the main argument will hinge not on the permission though, but rather if the use of code that is copyrighted in an AI model is transformative enough to fall under fair use. Obviously it's to be decided but I would imagine that because it wasn't a human transforming the code and/or hand selecting the code to put into the AI model, that it won't be considered transformative and therefore the use of the code doesn't fall under fair use.
I'm very curious how this case will play out.
To illustrate: GitHub could delete any project they want, and there would be no real recourse for the project's author. That is a service decision that they reserve the right to impose via their TOS. However, if they were to steal code from a user's private repository and violate the license therein, the author could sue for theft of intellectual property.
Again, the LICENSE file in the repo is not the only license for that code. A copyright holder can grant people licenses to their work with or without documentation and with or without that license being accompanied within their work itself.
By uploading code to GitHub, you are asserting that you can legally grant GitHub a license to that code for hosting as described below.
> If you're posting anything you did not create yourself or do not own the rights to, you agree that you are responsible for any Content you post; that you will only submit Content that you have the right to post; and that you will fully comply with any third party licenses relating to Content you post.
Note that this is literally only limited to the provisions set below; uploading to GH doesn't allow them to import or use your code in Windows or the Github codebase or anything like that, doing so would indeed be bound by the license terms you've granted the world via the repo's LICENSE file.
> 4. License Grant to Us We need the legal right to do things like host Your Content, publish it, and share it. You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time. This license includes the right to do things like copy it to our database and make backups; show it to you and other users; parse it into a search index or otherwise analyze it on our servers; share it with other users; and perform it, in case Your Content is something like music or video.
https://docs.github.com/en/site-policy/github-terms/github-t...
I think we're over here in our armchairs weirdly assuming that GitHub doesn't have any lawyers working for them. I think they know they're legally in the clear on CoPilot.
I'm not at all a lawyer, but in my opinion we observe that the non-automated version of AI-generated works (the act of making art and prose in the style of an existing copyright work based on the artist's observation of that work) is not illegal. The only thing that AI introduces is automation.
What I can't understand is people feel locked into Github because of the social features. To me they seem the least important part of Github, particularly with so many OSS projects running communities on Discord or Slack.
Hmm...Automated Inference? Automatic Infringement? Maybe we can make a nice backronym out of this.
IANAL, but I think Copilot is not a reasonable thing to include in these services.
Look, proving copyright infringement is downright trivial here, if copyright law applies. And that shows where GitHub’s defence will—must—lie.
(And for other readers unfamiliar with the parent comment’s phrasing: “end-run” is apparently an American sporting term which here makes “as an end-run around” mean “to circumvent” or “to work around”.)
DMCA was designed to catch not just normal pirates that share the copied content, but also "crackers" that figure out how to share content that is protected somehow, as such it offers various ways to infringe it without infringing the copyright itself.
Basically they are accusing MS of behaving like crackers, by removing stuff from code to allow it to get shared illegally.
Signing up for disperse litigation like this seems like a pretty ballsy move by GitHub, but hey, Microsoft presumably has in-house lawyers with lots of spare time.
a) at scale
and
b) make some other population happy
In other words, piracy should be perfectly fine, too, right? After all there are a huge number of users who benefit from it?
Agreeing to GitHub's terms doesn't try to assign copyright over your code, it grabs licence to use your code however they see fit which is¹ legally quite different.
Of course the real fun comes if someone agrees to their terms then uploads some of my code which they have to right to assign the licence to GitHub for. What come-back do I get in that case if I don't want my stuff used that way?
It seems odd to me that MS² who for many years strongly spoke against touching anything with the remotest whiff of GPL because of what it could legally do to your release requirements, are now more than happy to hoover up all the GPL covered code in GitHub and potentially mix it into their users' work output via copilot.
----
[1] in my not-at-all-legally-trained understanding
[2] current owners of GitHub, for those not paying attention
If it is shown that the license in the TOS is valid, the legal question might boil down to "is the TOS License broad enough to where nobody thought that it allowed their code to be used in for-profit ML models?"
user bluca works at microsoft, but i think their opinions are their own.
How does one check who owns a work if the work does not include the authorship information? (or if the work has been altered to have incorrect authorship information)
Edit: Wow, this is game changing. Markdown parsers need to implement superscript ascii character support!
Lowercase ⁽ᵃ⁾ Uppercase ⁽ᴬ⁾ Numbers ⁽⁹⁹⁾
To type them easily you'll usually need composition (sometimes called chording) support. Some Linux (and other Unix) distributions still have this built in by default, though last time I used Linux for much desktop use it seemed to be fading from common availability, otherwise you'll have to hunt for another method. On Windows I use http://wincompose.info/ (here [atlgr][^][1] produces “¹”, for instance, in the default settings) which is useful for a number of other things (I first started using it for accented characters like á on a UK keyboard). If you have a keyboard with programmable function keys then you could use its customisation tool to map some of them to produce the super-script (or sub-script, or other) characters you commonly want.
For less convenient typing, use your OS's Character Map or similar tool.
On Android, unless you have a different keyboard in use which doesn't support this of course, long press on the number on the touch keyboard gives superscripts as an option.
chucks iPad out the window
We only get the standard shift character as an option. E.g. 1 shows !, 2 shows @, etc.
I’d use superscripts all the time if it was on the keyboard. Anyone know if MacOS can do it? Other than pressing the weird globe key and searching.
https://apps.apple.com/us/app/unichar-unicode-keyboard/id880...
They are _not_ in ASCII. A few are available in some 8-bit code-pages that expand on ASCII's 7-bit character set, otherwise you need to be working in a Unicode-supporting environment (which is most these days, thankfully).
I disagree, IANAL, and I'm happy they are getting sued. The fact that they are are foremost a code hosting/collaboration company and the terms of service we all agreed to when creating our accounts was to have them host our code, and use it however they need in order to provide the service. The fact that they changed, post agreement, the service provided (from mere hosting/collaboration to feeding it into Copilot) should be an opt-in. I hope it's tested in court what the service is, because if you have a feature that (let's say) 1% of your users use, that's not the service, is it?
I'm pretty sure I didn't receive one about them using my public (although unpopular) open source code into their NN mixer.
Edit: Anyway a bit outside the point. It being, when your ever expanding set of services incorporate your ownership in ways unforseen when the agreement was made, opt-in would have been the agreeable aproach in my opinion. Even ignoring over the licensing woes, as that's something to be tested in courts with this lawsuit, and interesting to follow.
This isn't about privacy, it is about licensing (and possibly copyright). mhitza mentioned privacy as another policy, that you agree to upon sign-up like the terms of service, one for which updates are regularly announced.
> I'm also curious how you're certain your project was used?
Hasn't it been suggested that all public repositories at least could have been used? It makes sense to give the training pool as much information as possible.
The terms of service say you grant GitHub an implicit license to display your code. They also say:
"We may modify this agreement, but we will give you 30 days' notice of material changes."
Are you claiming that hasn't happened?
> Hasn't it been suggested that all public repositories at least could have been used? It makes sense to give the training pool as much information as possible.
Has it? I don't like to make assumptions.
> Are you claiming that hasn't happened?
Your post that I replied to explicitly stated that it hasn't.
Is that the case or was that one if the assumptions you don't like to make?
Not your code. Anyone's code that's uploaded to github by any third party. Under open source licenses, that's expressly permitted. However, it seems you're arguing that Github is not bound by the license under which they (and their users) acquired the code because of their TOS.
How many projects on github are put there by the original copyright holders? Perhaps it's more than 50%, but it certainly is less than 100%. So where is github's legal paperwork that shows that they're only processing the code that's copyrighted by the user who uploaded it, or that they received permission from third-party rights holders that did not agree to their TOS?
It's interesting, because Github certainly have the right to set whatever terms they like on their website. The dependency authors certainly have the right to set the licence terms on their code, and that gives me the right to vendor their work and include it in my upload to Github. But I, obviously, don't have the right to agree to Github's terms on behalf of the dependency authors.
I think the problem here is that Github assumes that everything I upload is my property, and that I have the ability to assign a licence to what I upload. This is not true for any project that vendors its dependencies.
Under other circumstances they don't need it. But if CoPilot is creating a derivative work including parts of that code without including the licence terms or attribution (as required by many licences) things are far more grey, or possible full black.
Some argue that the AI is unaware of the terms so can't be held responsible. Two possible counters for that: 1. it is the licence that gives you the right to use the copyrighted code, if you are unaware of the licence then why assume you have the righ tto use the code? 2. if I found some useful code that happened unbeknownst to me to be from MS, and used it in a way that I wasn't licensed to, and MS noticed, it is a pretty safe bet that they'd state ignorance of the copyright terms doesn't mean you can't be held to them.
Or another angle: the tool is allowing, even encouraging, people to use code or other materials in a way that infringes copyright (again: you don't have the right to use the code under most licences unless you give correct attribution and such) – the very conditions often stated as reasons for trying to ban other tools.
Plus of course the general argument: if this is entirely a non-issue, why is no Windows, Office, or SQL Server code in the training set? Surely they are great examples of how to do things to train the AI with?