Give up GitHub: The time has come
sfconservancy.org
sfconservancy.org
This is my favorite question about Copilot ever.
>there is nothing really surprising in this leak. Microsoft does not steal open-source code. Their older code is flaky, their modern code excellent. Their programmers are skilled and enthusiastic. Problems are generally due to a trade-off of current quality against vast hardware, software and backward compatibility.
[0] https://web.archive.org/web/20040401115821/http://www.kuro5h...
I can think of annoyances but awful? Come on.
Specifically, clear anti-user choices that exceed by far being "annoying":
* Making it exceedingly difficult or impossible to use the OS without logging in with a Microsoft account.
* Forcing the user in various ways to surrender data to Microsoft. Some of them can be disabled if you really go out of your way, others can't.
* Prompting me again and again to switch to Edge and other MS defaults. I've had the same install for a few years now and NO, I don't want to change to "Microsoft recommended defaults", no matter how many times you ask me.
* Showing the same "OS setup" screen after some updates, requiring me to pay very close attention to what I'm clicking, lest I select something MS is trying to lead me to. The amount of attention required from the user on those screens corresponds quite well with anti-user behavior.
This is hilarious. I recently got a new laptop that has window$ 11. After setting it up with a Non Microsoft email (which required some good fight), I tries to install some random app from the Microsoft store, but got a "something went wrong please try again" on the first screen.
It's pathetic. I haven't used Windows since Win 7 , which I basically installed for gaming. Seeing the latest version of the OS makes me feel sorry for them. That's why Apple with all their assholery is eating their lunch (on the flip side my wife just got a MBP m1 and I was pleasantly surprised that it has hdmi port, magsafe, several USBc ports. Apple seems going in the right direction.)
The telemetry makes this clear. Reboots and updates even more so.
The UI lag and stealing of focus ("oh, you're typing a document... too bad, I want to launch a new Explorer window that will immediately steal focus") make it clear that the computer is in charge and will probably listen to your requests, but on the timeline it chooses.
Past MS engineers have been commenting for a decade on how MS has grown too big, can't manage, and has become a monolith "too big to fail". By nature when engineers are small pieces of a giant machine, they don't do their best work. And those with the experience move on to better things.
I'm pretty sure there are a lot of those.
This may be the dumbest move from M$ that I have read on this thread! Sure, companies need to protect their private IP, but this really feels like creating unnecessary friction for no good reason...
That's definitely something that most large corporations do.
Because it probably mines one bitcoin block every time you click on something. No sane codebase could possibly be so abysmally slow.
OpenAI Codex (which copilot grew out of IIRC), Amazon and Salesforce versions of Copilot exist. Huggingface Bloom was trained on a sizeable amount of public code. Tab9, now behind, was one of the earliest to combine public code repositories with Deep learning for smarter autocomplete. The data requirements for Transformer scaling mean any and all public facing repositories will be assimilated, whether Github, Gitlab, Stackoverflow or so on.
Wish more energy was spent on how to fund pretrained models that will also run efficiently on CPUs, fine-tuneable to one's language and local environment. Removing reliance on cloud services.
Curious about people's opinions on Dall-E 2 or Google Image-gen, which parallel pretty much the same thing with Renders, Illustrations and Paintings, or upcoming models doing the same for voice acting and music. Coders seem more excited about the potential of those tools.
CoPilot is being offered for widespread commercial use, so it's held to a higher standard. Respecting copyright is much more important when you're building a business and not just sharing fun AI art on social media.
But that's not as damning as it sounds.
First, we know Copilot, if given the right prompt and told to autocomplete repeatedly without any manual input, can regurgitate bits of code seen many times in many different repositories, like the famous Quake fast inverse square root function and the text of licenses. That doesn't mean it does so under normal prompts and normal use. Perhaps it does sometimes, and that would be a real concern. But any regurgitation that isn't under normal use, which only happens if the user is trying to make Copilot regurgitate, is not a problem when it comes to copyright violations of open source code (since anyone trying to violate an open source license can do so much more easily without using Copilot), yet it may still be a problem when it comes to leaking confidential information.
Second, whether something is a copyright violation and whether it risks leaking confidential information are somewhat orthogonal. A copyright violation usually requires at least several lines of code, and more if the copying is not verbatim, or if the code is just a series of function calls which must be written near-verbatim in order to use an API. On the other hand, `const char PRIVATE_KEY[] = ` could hypothetically complete to something dangerous in just one line of code. That said, it almost certainly wouldn't, since even if a private key was stored in source code in the first place (obviously it shouldn't be), it probably wouldn't be repeated enough to be memorized by the model. Yet…
…third, the risk tolerances are different. If, to use completely made-up numbers, 0.1% of Copilot users commit minor copyright violations and 0.001% commit major ones, that's probably not a big deal considering how many copyright violations are committed by hand – sometimes intentionally, mostly unintentionally. (When it comes to unintentional ones, consider: Did you know that if you copy snippets from Stack Overflow, you're supposed to include attribution even in any binary packages you distribute, and also the resulting code is incompatible with several versions of the GPL? Did you know that if you distribute binaries of code written in Rust, you need to include a copy of the standard library's license?) But when it comes to leaking confidential information, even one user getting it would be somewhat bad (though admittedly Microsoft does distribute much of their source code privately to some parties), and taking even a small risk would be a questionable decision when there is a ready alternative.
If Microsoft/Github ever made that argument, that also means that when Copilot is using GPL software as input, the output can only be released under the GPL.
Copyright licenses don't apply to small snippets, no matter if you think they do, and learning and applying other people's code isn't prohibited by the license, and thank god, can't be prohibited.
There's also a setting at https://github.com/settings/copilot (link only works if you've signed up for copilot) that will check any suggestion on the server against hashes of the training set, and block anything that exactly duplicates code in the training set (with a minimum length, so very common code doesn't get completely blocked). Users must choose the value for this setting when they sign up for copilot.
source: I work on copilot at github
Just because they use FLOSS licenses, does not allow them to evade things like Affero GPL3. And, to that end, if they are using Affero, I want the source to the whole copilot infrastructure -or- proof they used no AGPL3 code anywhere.
A better question would be whether they would take legal action against a competitor that creates a copilot equivalent and publicly states that they trained it on leaked, proprietary M$ source code. That would actually be an example of hypocrisy.
Because these models work better with more data and presumably this a lot of high quality data that they already have lying around anyway? Because there no downside according to their own reasoning? Because it would shut up a lot of these criticisms right away? Because marketing would be so much easier with that kind of dogfooding?
In short: because according to their own story there would be only upsides, no downsides.
I also want to know why people think their code is so special that no one else could have ever come up with it independently. Each and every opponent of Copilot is the best developer ever, I guess?
That said, I don't understand the choice to use GPL for any reason, so maybe I'm not equipped to understand the arguments against Copilot. Forcing your code to be open forever isn't freedom, it's the omission of freedom. Someone using your (for example) MIT-licensed code in a closed-source commercial software project doesn't "un-free" the code you released; your code is still exactly as open and as available as it was before, and zero freedoms were lost by anyone.
Which is a freer society: one that restricts late night partiers from playing loud music in residential areas, or the one that does not?
Determining what freedom should mean is not, and has never been, a simple matter of "well, if you make any restrictions on it, then it's not real freedom, so everyone just gets to be free!" It's all about finding balance, and dealing with nuance, and all that frustrating hard stuff.
No, that’s forcing restriction on all users of your code.
Usage restriction is the opposite of freedom…
Forcing all of your code to be GPL is like saying “I am on a diet, so now I will force everyone else be on the same diet. Freedom!”
By analogy, there is a law against me putting handcuffs on another, and in fact the police would stop me from doing so. Did the police protect freedom? Aren't they restricting me from handcuffing others?
In a similar manner, under the MIT I can restrict my users from modifying and compiling my source code. Is a license that means I have to let my users modify code restricting freedom? Isn't it ensuring freedom of others, in the same way that making laws of "you shall not handcuff others for no reason" is ensuring freedom of others?
GPL license ensures that your users will keep the same freedoms that you got.
Of course, there's an inherent conflict - the freedom to oppress others is incompatible with freedom from oppression.
Nobody is forcing anyone to use the code.
If they chose to use it they have to abide by the licensing terms because that’s how it works. If the people laboring for free to produce this code don’t want it to be used in a proprietary application then tough luck, write the code yourself.
Every time the GPL comes up someone drags out this same old dead horse to beat on a little bit more.
until the time comes when a tax department gets the funny idea to use it, and forced you to use it, or people with guns come to your door and haul you away in the morning.
(edit: formatting)
This is a terrible analogy. Here’s a better one: I’m holding a potluck. If you decide to come, you can eat all you want. If you take food from my event, you can’t hoard it, you must share it, even if you’ve “made it better” by changing it somehow after you left.
Don’t like my rules? OK, don’t come to my potluck.
Suppose that there's a law that states that water and access to it is always supposed to remain public, because water is a public good.
Suppose that someone comes tomorrow and starts claiming ownership of all the water springs in your country, he becomes the only entry point to get water, and you have to pay him a fee every time you open a tap.
Is he still free to do so? In other words, is the freedom of someone who restrict the freedoms for everyone else still a form of freedom that is worth even considering, let alone respecting?
Because the foundation of your ideas is exactly the reason why capitalism fucked things up and just let a bunch of jerks get rich without merit.
Boost Software License - Version 1.0 - August 17th, 2003
Permission is hereby granted, free of charge, to any person or organization obtaining a copy of the software and accompanying documentation covered by this license (the "Software") to use, reproduce, display, distribute, execute, and transmit the Software, and to prepare derivative works of the Software, and to permit third-parties to whom the Software is furnished to do so, all subject to the following:
The copyright notices in the Software and this entire statement, including the above license grant, this restriction and the following disclaimer, must be included in all copies of the Software, in whole or in part, and all derivative works of the Software, unless such copies or derivative works are solely in the form of machine-executable object code generated by a source language processor.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE, TITLE AND NON-INFRINGEMENT. IN NO EVENT SHALL THE COPYRIGHT HOLDERS OR ANYONE DISTRIBUTING THE SOFTWARE BE LIABLE FOR ANY DAMAGES OR OTHER LIABILITY, WHETHER IN CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.
Permission to use, copy, modify, and/or distribute this software for any
purpose with or without fee is hereby granted.
THE SOFTWARE IS PROVIDED "AS IS" AND THE AUTHOR DISCLAIMS ALL WARRANTIES WITH
REGARD TO THIS SOFTWARE INCLUDING ALL IMPLIED WARRANTIES OF MERCHANTABILITY
AND FITNESS. IN NO EVENT SHALL THE AUTHOR BE LIABLE FOR ANY SPECIAL, DIRECT,
INDIRECT, OR CONSEQUENTIAL DAMAGES OR ANY DAMAGES WHATSOEVER RESULTING FROM
LOSS OF USE, DATA OR PROFITS, WHETHER IN AN ACTION OF CONTRACT, NEGLIGENCE OR
OTHER TORTIOUS ACTION, ARISING OUT OF OR IN CONNECTION WITH THE USE OR
PERFORMANCE OF THIS SOFTWARE.Please feel free to use my code in any way that its license permits: attribution for the permissive licenses, share-and-share-alike for the copyleft licenses. Those license terms are the price of the code, no different from a proprietary product's "this costs $x" or "this costs $x/month". I'm happy to give away most of what I work on every day, and I ask that people 1) give credit, and 2) in some cases, share under the same terms, and 3) in many cases, don't sue me or other users of code I've written over software patents (which shouldn't exist).
If the day comes that copyright goes away, and we can freely copy and share the code of any currently proprietary software and other works, I'd celebrate that. Until then, I don't want an asymmetric situation in which proprietary licenses must be adhered to but Open Source licenses are ignored.
a) Those who do have the source code to share it. Sometimes the source code is available but it can still not be freely used and/or shared.
b) Allow modifications and redistribution of the binary artefacts, which for is sufficient for many goals.
c) Remove all concerns with reverse engineering and allow e.g. decompiling programs and sharing that source.
Also, remember that the GPL already does not make source available externally if the modfications are only used internally.
I'm worried about exactly the opposite: having Copilot help me write code that seems quite generic to me, but which in fact makes my code subject to a license I don't even know about, and/or simply violates copyright.
For an open-source project this could be embarrassing but probably fixable. It gets more complicated if FAANG is doing due diligence on your company. I can see Copilot being both an accelerant and, later, a liability for startups.
I guess time will tell how much acquiring companies (my worry) care about Copilot. Given the difficulty hiring good devs, and the productivity level of body-shop devs, I see it getting a whole lot of use very soon, acknowledged or not.
But also.. In that case, when I commission an artist to paint my portrait, surely I can't claim to be the artist.. But I'm no lawyer.
I'm not sure there is a contractual agreement in GitHub's co-pilot that says: "Any code you write here is commissioned work". But honestly I didn't read the T&C's.
So I think you MAY have debunked my analogy, but not the main reason for the analogy.
As software takes a back seat (or at least a "normal" seat) in society, would we see a normalization of income? Could this be hastened by the development and introduction of tools such as copilot?
Potentially, unless there are new / better things that humans can claim they can provide compared to AI tools. This is the point where I think you and I agree, and I think it's your primary argument in any case (unless I'm mistaken).
Compare the visible output of someone writing in assembly vs someone writing on top of a modern web framework. Is assembly harder? Yeah. But the web framework is going to give you a usable product in a fraction of the time with way more features. And that's worth more money to the company you work for.
It's always going to be a knowledge worker's job. It's always going to reward experience and creativity and attention to detail. A lot of programming is looking at the world, seeing a gap in what exists, and figuring out what best fits that gap. An AI can't do that. Programming is making 1000 tiny decisions that can't possibly be specified completely by a product manager and need a human to weigh the tradeoffs.
Thats what everybody in the chess world said: "AI can decide low level stuff. This one move. This small attack on a rook. What it can't do is conceive of how to take a bunch of different tactics and put them together to produce a game of chess."
...Until Deep Blue beat Garry Kasparov.
> It can't tell you if you should use postges or mongo.
Yeah, and then came: "It may be able to play chess, but it can't tell you how to play Go."
Look how that went.
So, unless you are a code monkey punching code into autogenerated skaffolding all day, your job is safe.
This isn't true. As an analogy, consider that forcing people to not own slaves isn't the omission of freedom. See also https://www.gnu.org/philosophy/freedom-or-power.en.html
Like thinking you're ending slavery by freeing all the current slaves but not making it illegal to own, buy, and sell slaves, or capture previously free people into slavery. Guess if you'd have slavery again very soon?
The analogy is about freedom vs lack thereof, not manual labour vs software. And as you see, it works very well.
Truly a classic.
Really? What exactly does this CoPilot thing actually spit out? I can't help but think that it spits out near verbatim, which in the UK is probably dodgy on Copywrite.
You then go on to decide that the GPL isn't for you. That's fine. You even explain that you are ill-equipped for something. That too is fine.
You are not a fan of free or "libre" stuff. That comes across loud and clear. Thank you.
Nobody has claimed that they want this. People just want derived work to adhere to the license they chose for their project.
>I also want to know why people think their code is so special that no one else could have ever come up with it independently. Each and every opponent of Copilot is the best developer ever, I guess?
Would you feel the same way about ripping off game assets, or music?
I think you just have an axe to grind with free software in general based on your messages and the general tone. Just because you don't understand it doesn't mean that the ideas are invalid.
I am also curious why copyright laws should protect proprietary software, music, games, writing, etc but not apply to my software, even if it isn't the highest quality work?
Also, there's no way for anyone to know what portion of code that I commit was hand written vs. generated, so you kind of have to treat it all as written by the committer anyway.
Though this does bring up interesting questions about what happens with things like automated PRs that fix bugs / update dependencies... are those then non-copyrightable? ¯\_(ツ)_/¯
Just as much as a hero riding off into the sunset is not copyrightable in a movie script. However, a hero riding off into the sunset with bananas in the pistol holsters would be.
This is what I would want to hear more about when discussing if Copilot violates copyright.
This has some interesting implications – for example, it means I can't mirror somebody else's (open source) code on GitHub without their explicit agreement.
So any code uploaded by someone other than the copyright holder renders someone liable to be sued for copyright infringement, AFAICS. The only question is whom it makes liable -- the uploader, GitHub (=Microsoft!), or both?
I can see arguments either way: The uploader is clearly infringing by giving away a right that isn't theirs to give. But so is GitHub / Microsoft, for using a "right" they haven't been properly given. So I'm provisionally leaning towards "both".
> I can't mirror somebody else's (open source) code on GitHub without their explicit agreement.
Who is doing the "mirroring" -- you, in uploading the code, or GitHub / Microsoft in actually hosting it, keeping it available for download from their "mirror"[1] site?
___
[1]: Is that even the correct terminology nowadays, when AIUI for lots of projects GitHub is their primary code repository?
either code is owned by its licensors or it isn’t.
> I also want to know why people think their code is so special that no one else could have ever come up with it independently.
I’ve never heard anyone argue this in the real world, ever —and i’ve been involved in this space for years.
If someone doesn’t want our code, then they can go ahead and write their own from scratch. We’re certainly not stopping them.
many people do seem upset at us that we’re sharing code, tho. particularly that group who primarily make their fortunes from other people’s work.
GPL uses a different definition of freedom, which I prefer. They look at consequences of restrictions / permissions, and their implication on freedom (not just for me, but for everyone). So some restrictions can lead to actually more freedom, while some permissions can actually decrease freedom.
This is similar to gun-control. While it reduces freedom for gun owners, it allows everyone to be more free of hanging out anywhere they want without being afraid of being shot. Similar arguments can be made for vaccine mandates.
So GPL restricts usage of software because in the long term it gives back power to users, which will be more free.
Eh. I see what you’re saying about gun control, but the idea that “some restrictions can lead to actually more freedom, while some permissions can actually decrease freedom” is actually very American.
The free software movement says that everyone deserves software freedom. The Declaration of Independence similarly says “We hold these truths to be self-evident, that all men are created equal, that they are endowed by their Creator with certain unalienable Rights, that among these are Life, Liberty and the pursuit of Happiness.” While I haven’t found a source confirming it, I think that the founders believed that the freedom of speech was one of these unalienable rights.
The GPL puts restrictions in place to make sure that downstream projects give users software freedom. The Constitution put restrictions in place to ensure that the federal (and nowadays the entire) government doesn’t interfere with our unalienable rights.
Take a look at how the first amendment is worded:
“Congress shall make no law respecting an establishment of religion, or prohibiting the free exercise thereof; or abridging the freedom of speech, or of the press; or the right of the people peaceably to assemble, and to petition the Government for a redress of grievances.”
The first amendment does not grant the freedom of speech because it doesn’t need to be granted. From the founders’ perspective, god already grants the freedom of speech to everyone forever. The key phrase here is “Congress shall make no law”. The first amendment is restricting Congress to ensure freedom.
The idea that “some permissions can actually decrease freedom” is also present in the Constitution. For example, take a look at Article I sections 8 and 9. The framers of the Constitution could have given Congress the power to pass any law. Instead, they chose to specifically enumerate what Congress can and cannot do.
Perhaps, though, most Americans don’t know much about our founding and think that freedom=just let me do what I want. I don’t know.
It’s the users of that closed-source commercial software that lose freedoms.
How many times does it have to be stated? GPL is for the users.
For this reason, I had idea to make up a new license (although I will not write most of my ideas here but will do so elsewhere). But, its main working would be: mostly you can do whatever you want (including omitting attribution and copyright notices) without worrying about the license, but you cannot use legal processes (such as lawsuits, DMCA, etc) to prohibit these freedoms to any downstream recipients (regardless of how many). The license would also ensure patents can be used freely, disclaimer of warranty (if the license is included in the copy and the recipient has not paid for the copy), and some other things to ensure freedom (although there can be some restrictions on the use of trademarks (e.g. to avoid false advertising), and some things to avoid working around the freedoms in certain ways). You can be forgiven any number of times, though; the license will not be terminated. Furthermore, for a practical reason of license compatibility, relicensing by GPL3 and AGPL3 (and possibly also CC-BY-SA 4.0, for works other than computer programs) are also allowed, as long as you have a copy of the source code and can satisfy the terms of those licenses.
What a disingenous reply. FOSS licenses do not grant ability to replicate "in any way" that you wish. You still have to comply with the license terms. What the hell is wrong with you?
> I don't understand the choice to use GPL for any reason ...
The reason is people like YOU.
A similar line of thought is the "paradox of tolerance", which posits that if a society tolerates the intolerant, the tolerance of that society will lessen.
Do you feel the same about other creative processes as well? Can I rip a Justin Bieber's song and say that it's mine just because it's a shitty song anyway, so who cares? Or does this only apply to software because software is somehow an "inferior" art? Do licenses even have any legal value to you?
On the contrary, forcing your code to be open forever is the only way to preserve freedom.
> I don't understand the choice to use GPL for any reason
No, you just don't understand the GPL.
Or, for some reason, pretend not to.
While GitHub might have a license to use that code to train the model, it’s debatable what license applies to the output of the model, and what users of the model can do with it.
It’s possible for an AI to reproduce something so close to the original that it would be considered an infringement on the original work.
Co-pilot has issues, ergo github is going the way of sourceforge, and so now we must abandon github? Do I have that reasoning correct?
We need to:
- migrate the bug queue
- have all links in commit history break
- application integration with githib for bug reporting be removed
- update documentation
- find a new website host (and no longer github.io)
- find a new CI/CD (we were already burned by travis, github workflows are nice)
- teach our user contributors to actually use git! There has been a lot of heartache from them that they have to use the pencil icon on a web ui to edit config files, now we have to take them back to using a git GUI client! We were on github before that blessed pencil icon feature came out, there was no end to the wailing about how unapproachable the process was (super frustrating when users see they have to do something.. frustrating for us because our users wanted to just email is stuff so we could then do the uploading to git work)
- lose all PR history
- migrate project tracking
- find a new place to host release artifacts
- update our website to use a new distribution URL (the website scrapes github api to get latest version for download link; it's nice never updating website as we do releases on every merge)
- figure out and migrate repository permissions. (We have a hundred repositories of user generated plugin content, everything about migrating that would be a lot of work and missing important features)
- lose our search ranking and rebuild our SEO
What else to add to this pile.. and all because co-pilot smells?? Meanwhile all of that work is busy work, and not at all feature-pare. That kind of migration would take a long time, seems like that pivot without good reason is the worst kind of churn. Convince me this article is not a temper tantrum about copilot...
The money model, as is for many free to FOSS tools is that by getting devs tooo use those tools, they'll carry forward to their professional lives and recommend the adoption by their companies. That does happen in practice, so it seems like the money model will not necessarily flip like it did for sourceforge (which kinda was garbage and the only game in town)
I would disagree about the industry dependency compared to FOSS. Many companies are not on github
So, that is to say the dependency aspect is a concern. So far Microsoft has overall been a steward for FOSS and copilot is not at all nearly enough to lose that trust. It is always a bit nerve wracking to place your balls in someone else's hands... and it was concerning when MSFT bought github.. but they have not been evil, not even close yet (in the grand picture)
I'm curious if anyone can find references, though when I researched market share of code hosting companies a few years ago (for a private company that was moving off of BitBucket), it turned out that there were more private companies on Gitlab than Github. Github though had a big advantage for hosting FOSS. We wound up moving to Microsoft Azure because the scrum boards and Microsoft integration were appealing and familiar to the company. I don't see it being analagous as Windows desktop control in the early 2000's.
In any case... why github? What is so unique about it that you can't even consider other possibilities? I guess soon people will be non-ironically saying "no one ever got fired for hosting their code at Github" and turn a blind eye to perfectly usable, open alternatives who does not lock us in.
For what? Fear of taking responsibility for maintaining the basic tools for their job? If that is the case, you can always pay for other smaller, independent companies who can host at competitive rates.
Anyway, you do you. I'm tired of playing Cassandra, and I'm tired of seeing people giving in to convenience and general conformity.
So you provide the system, it hosts itself.
disclosure: I work for github
> barrier of entry to contribute to your code
What barrier, may I ask ? > it's a lot harder for others to discover it.
The vast majority of stars I got was from a HN submission, not from organic discovery through GH search.Even low barriers of entry can cause a big drop in user engagement.
(I will not get in a tangent about web3, but that is the one thing that web3 skeptics always fail to acknowledge is how the current web is broken in that regard. We were promised open protocols, and we end up with a handful of companies building their own walled gardens)
The only way that Github would get any modicum of credibility would be if they joined the effort from codeberg/forgefed and integrated with activitypub. As it is now, github will be nothing but a mirror for my repositories that I will be hosting on gitlab and/or my own gitea.
No. Stopping believing this shit. It has never been easier to not depend on any of them.
There are good, perfectly usable phones with de-googled Android. You can buy laptops and desktops that run Linux without issues out of the box for years. You can even game on Linux today better than you can on an Apple box. You have legions of people working on different open source projects that make hosting your own server a matter of point-and-clicking.
It's not an Sisyphean job to keep yourself away from bad tech products. It takes some discipline, but so does anything worth doing.
And if you really just want to pay to get rid of any "headache", why not then pay for an open source alternative so you can be safe knowing you won't get locked in?
If GitHub didn't provide value, then it wouldn't be where it's at. Considering that collaborating on FOSS before GitHub was a mess if you weren't technical, I'd say GH has earned their spot.
Also, there's no way to avoid centralization. At some point somewhere, you're relying a mega-corp for critical services, and if not you, someone else.
The way I read the article was that by being on GitHub, you are implicitly agreeing to no longer be a FOSS project as regards licensing. GitHub customers can use Copilot to generate proprietary code that's identical to your project's code (several articles I have read call this overall idea "laundering through Copilot", which sounds incendiary but accurate to me) without needing to respect your license.
The other stuff you said is...kind of irrelevant. Sure, you get a lot of convenience from GitHub. If you don't care about software freedoms in the libre/copyleft sense and regard a "do whatever you want" license as the best, then it's probably fine to keep using GitHub.
In this case, they have violated our license. The license states how it cannot be changed, and implicit changes by the hosting company is not one of them. Isn't there a legal case here? If no, then the license is not violated. This makes me wonder why not take Github to court vs a tamper tantrum (we're taking our marbles and going home!)
> the other stuff you said is...kind of irrelevant.
It goes to show that Github is providing really, an excellent service to FOSS. We could migrate to GitLab, though, why?
I know that sounds like 'fan-boy', but the list is large, useful, and all highly available and free to FOSS. By those measure, it's a good service.
> Sure, you get a lot of convenience from GitHub.
The items are more than convenience, they are core to our application and project. Uprooting them is not a small task.
The automatic integration with issues is excellent for example, before we did that - we had no idea how many users were seeing errors. We run a thick client that is downloaded and added an integration to upload error reports to github issues. That has been pretty invaluable. So, we have to move all that, to another host: who is to say that other host will always be a better FOSS steward? who is to say that other host has anyone near the level of features? An integrated CI/CD and hosting of release artifacts is huge.
So, on the premise that our license has been violated, instead of sue'ing, we should take our FOSS somewhere else? Again, the list is large, we need to migrate all of that. It took a long time to get out of sourceforge, it's not even more work to get out of github because we automated so much (to allow our team to scale better). It's not just a matter of 'convenience' to go somewhere that is feature-sub-par and spend the better part of a year to do that and nothing else.
[edits] Clarity, conciseness
Mainly because it might be more in line with traditional FOSS ideals to take our marbles and go home? It's hardly a temper tantrum. After the Linux kernel got into one too many arguments with Bitkeeper, there was an inflection point where Torvalds got fed up and just wrote git instead. I think this is an inflection point like that. And "taking Github to court" is much easier said than done for a FOSS project, although I assume that someone will be doing that at some point.
> So, we have to move all that, to another host: who is to say that other host will always be a better FOSS steward? who is to say that other host has anyone near the level of features? An integrated CI/CD and hosting of release artifacts is huge.
In addition to GitLab, I think Sourcehut offers similar features and is also FOSS (AGPLv3).
> So, on the premise that our license has been violated, instead of sue'ing, we should take our FOSS somewhere else? Again, the list is large, we need to migrate all of that.
I don't think anyone is telling you what to do, just offering suggestions. My suggestion is that switching to a forge that respects FOSS is a reasonable alternative to legal action against GitHub or just ignoring the potential license violations via Copilot. You can of course choose any of the alternatives, or the status quo; which is what you seem to have largely convinced yourself is fine. Good luck, and thanks for working on FOSS anyway!
Is copilot really a violation of GPLv3 (our license)? How is co-pilot different from someone lifting sections of code? Lifting a few sections of the code is a world different compared to re-distributing the entirety of the source code, or forking the project and replacing 2 or 3 letters from our brand name & then redistributing that.
I think the article needed to go into more detail about how that really is a violation of a license. This seems like a similar argument that was made in court whether the Java APIs themselves could be copyrighted. Can an algorithm be copyrighted or licensed? If someone uses the same algorithm as found on FOSS, have they violated the license of that FOSS?
Then the reaction, instead of pursing litigation, and/or communicating and working with github, the reaction is we should 'cancel' github and move to.. gitlab? Is that even an answer? If we think algorithms are copyright-able, wouldn't have any kind of code search be a violation? Would allowing for any kind of transcription of code be a violation? If so, then seemingly having the source be open would invite this.
I think this gets to the heart of FOSS in some ways. It's closed for privatization, open to the community, and what matters is the software provided. If someone cribs the project to configure a Feign client, or set up a unit tests with DbRider - it's okay! It's the same thing as viewing HTML source code, learning a cool javascript trick by looking at how some website did that trick - is part of the openness.
I wonder then, is the point of FOSS openness only to allow others strictly to view and edit the code for the purpose of contributing back to that exact software product? Or is the openness more than that, and that others are going to use the software in creative and novel ways, and use it for learning and who-knows how else (all pursuant to GPLv3).
Yes. Many projects will see your license and never copy your code if their licenses is incompatible with yours. Copilot disregards your license and offers your code or derivations of it to anyone who asks the correct questions, breaching your license both in ethical and legal level.
You can't lift a function from a GPL licensed codebase and plant it to a MIT licensed codebase. It's that simple.
Which I think creates an interesting debate, where is the line of general technology vs the intent of not allowing private companies to re-package FOSS for their own benefit?
> You can't lift a function from a GPL licensed codebase and plant it to a MIT licensed codebase. It's that simple.
If that function is a 'pad-left' type of function, is there an ethical violation? Or is this similar to learning Javascript by viewing the source code of webpages? At some point, functions are hardly unique in what they do, and there are only so many different ways to write a 'pad-left' function. I mention this question not to refute what you've said (I think I agree there is a likely license violation), but to explore where the line is for the ethics. I mean, an inferred implication could be that software developers stop looking at HTML source in order to learn. If you learn how to write a standard algorithm or function, and then you reproduce that later in private software, that is "lifting code". It's not much different committing that to memory then doing an outright copy-paste of a 2 or 3 liner. This makes me also think to the variety of shell scripts and standard bash'isms, if a single FOSS project uses a 'sort | uniq -c | sort -n', does that mean a MIT codebase has to re-invent a new way to do that?
It's an interesting debate for sure. When a program is divided into many functions, and when there are many utility functions, the debate indeed gets blurry. Moreover, many of the open projects in today's world are web applications and "simple" in nature.
In this context, simple means that the application can be composed via many simple, relatively general functions with a unique asset set and unique way of connecting them, so the magic (or sauce if you pardon the term) gets more abstract.
On the other hand, we have another group of programs, which can be similarly called "complex" programs. The biggest difference is these programs use complex, one of a kind functions.
Consider Blender, KiCAD, many open source and GPL licensed scientific software, libraries like Eigen, video/audio encoders, CD/DVD tools. Even a three liner in these programs and libraries can be a game changer.
I have a research oriented code, and a ~25 line function in this is worthy of its paper. I published a paper, and didn't obfuscate the language and algorithm, so you can implement it if you want.
However, if I open the code of this algorithm with AGPL3, can you say it's too small to be licensed? I don't think so. A more general algorithm can be ruled out as "too simple", but "fast_inverse_square_root" or my function or any other math heavy secret sauce can't be excluded because it's 2-3 lines.
As a result, while this issue needs serious discussion, we need to understand that even a small piece of code can carry a lot of research, knowledge and advantage in itself, regardless of its size.
So, while lifting a left pad from a GPL code can be understandable up to a certain point, getting the secret sauce, improving it and keeping it closed or merging into a similar, but incompatibly licensed software is inexcusable.
At the end of the day, this means we need to defend our GPL codebases, because we can't protect our secret sauce functions without defending the simpler ones.
If it's reproducing vebratim your code, yes it's a violation. And this is no different from others copying sections of your code.
The license that you chose says that people are allowed to do many things (e.g. copy and modify the code), provided they also fulfil some obligations. If they don't do that, they are violating the license.
Please note that this is not specific to GPLv3 (your chosen license). Other licenses, BSD-3-Clause, for example, also give rights (e.g. to copy and modify) and have their own set of obligations to be fulfilled (e.g. "Redistributions of source code must retain the above copyright notice, this list of conditions and the following disclaimer"). If someone redistributes source code (not the "complete" source code, even a "section of code" to use your words), they have to fulfill the obligations, or they violate the license.
I get the impression that this article and it's reaction is largely a lot of people wanting to cry big brother and mostly just shit on Github & Git. IMO it would be more appropriate to look for solutions and post a public letter to github and/or just pursue actual litigation before advocating FOSS projects spend inordinate amount of effort across the board to vacate Github.
You are correct that the problem is hard, because adding licensing info to training data and subsequently use it is something not yet accomplished by anyone (who at least has spoken publicly about it). It might be the only way out would be to have different training sets according to license, but then you lose the advantage of scale.
On the other hand, I believe it's a little much to ask SFC to do the cutting-edge research and propose solutions to Microsoft. Let's not forget that, when faced with the problem, their chosen path was "completely disregard licenses (for now), ask legal later". Many legal opinions are that there would not be any issue if copies of the input code were not reproduced verbatim; unfortunately this is not the case.
Looking forward to more exciting developments!
I think the issue with Java's API was actually terribly explained even by Oracle's lawyers, Google stole that API, first because it already had a large amount of developers familiar with it so they could benefit of cheap code monkeys developing for their platform, added modifications (memory management) which under Java's open source licence (Oracle owns java yet it is still open source) they should have made public for everyone's benefit (in a true Open Source spirit), modified the packaging mechanism (slightly, you know .apk), so instead of you know having a package (jar) that can run on a desktop or rather on any JVM it could only run on the devices with their* OS (or rather dalvik implementation) and guess what, now that there are some options to run those on other environments the "casually" decided to change it again instead of letting people benefit from the rich APK ecosystem somewhere else than Android
Of which none of your examples are in any way applicable to the case of google violating the Java copyright.
The Java Specification gives terms and conditions for usage which is what the lawsuit was about and not some laundry list of things which annoys a random internet person. In fact, not annoying your users isn’t even mentioned once in the conditions to call an implementation “Java” strangely enough.
Or I’m completely wrong and oracle doesn’t spend enough on their legal team.
Can you point me to those articles? I have seen the Quake thing, but "your project's code" is not like that one function in Quake (namely, it's not duplicated hundreds to thousands of times through countless license violations already).
For people that care to respect copyright, there's a copilot setting to block exact copies of code in the training set (which only happens a tiny percentage of the time, unless you're actually trying to make it happen).
For people that don't care to respect copyright, git clone is a way more efficient way to violate your license.
How is that different/new? Anyone has always been able to just take your code and put it into their proprietary projects...
Forget Copilot. Even without that you’ve put all of those services in one centralized basket, fully controlled by a for-profit company (with proven track record of unethical behavior).
And not only that, but this is true for the vast majority of FOSS projects!
I touched on it, but I still see this for-profit model as being compatible with FOSS. Specifically, make it free to FOSS so that those professional software developers then move to adopt the same platform at their private companies. There are lots of examples of for-profit companies that provide free-to-foss platforms.
For example, just considering code scanners, Codacy, CodeClimate, Snyk, LGTM, all used by this project and all have the same for-profit model (but free to FOSS). We also use 'install4j', which has the same free-for-foss model and is a for-profit company.
I think the heart of it is the argument that MSFT is going to expand that for-profit model and leverage data in a way that violates FOSS licenses? Is copilot an example of that? Should this all be taken to the extreme and declare that MSFT is the evil empire for this and by extension so is Github?
(a) since it is proprietary, you can't see where their hands are
(b) the fact that they haven't done anything also means you have no signal of their intent once they do. Not doing anything would be the logical play for a period of time whether you had evil intent or not.
I don't understand why you're essentially resetting their reputation in regards to GitHub.
IIRC Windows XP was the first version with online activation checks aka "Windows Genuine Advantage" so I think you need to go further back.
The question is not whether Microsoft should be put into the evil empire category, but if they really left it or are just pretending like they did every time before.
This is only a serious problem if you are not paying them for it. The current situation is how 99% of thr world works everywhere.
You are absolutely right that this all is a huge pain in the ass. GitHub, and later Microsoft, played their cards well. The product both works well, and also creates such a moat of vendor lock-in that it won't make sense to leave.
If the article's writer read this, please add this as one of the reason why FOSS need to move from GitHub.
Obviously.
> Co-pilot has issues, ergo github is going the way of sourceforge, and so now we must abandon github? Do I have that reasoning correct?
No, you don't. If you had actually read the whole article[1], you would have noticed that they also listed several other issues which far predate Copilot.
___
[1]: You need to correct this part of the guidelines, dang. Some times -- like here -- the injunction against saying this is just fucking wrong.
If this is just fucking wrong, I wonder about the implication. Would you say, I am just fucking right? I would be careful whenever having any such conviction. If I were already a hater of github, this article would have resonated with me a lot more. Perhaps there is some kool-aid drinking happening here? Maybe not and everything is totally reason to abandon Github as a hosting platform.
> you would have noticed that they also listed several other issues which far predate Copilot.
I don't quite see that list. I do see this article as quite focused on co-pilot. Though, I do see this:
> There are so many good reasons to give up on GitHub, and we list the major ones on our Give Up On GitHub site. We were already considering this action ourselves for some time, but last week's event showed that this action is overdue.
Which links to this https://sfconservancy.org/GiveUpGitHub/ (I'll point out here, linking to a list is different from listing the other points, so.. 'fucking right?')
Re-capping that list from sfconservancy: (1) co-pilot (2) contracts with ice (3) githubs hosting code itself is not FOSS (4) no self-hosting with github code options (again, github hosting is itself not FOSS) (5) work to discredit copyleft (6) wholly owned by MSFT
In this article, it essentially says that co-pilot was the last straw, and a decisively large one at that. So my response is still, this is the last straw where we need to abandon github?
If you are fully vested in the other reasons, then I think that this article would be preaching to the choir.
For me, points 1-6 are bothersome, but still just 2s and 4s on the 1-10 scale of fire alarms. My personal take on this list:
(1) co-pilot: seems problematic, perhaps github can fix it. Maybe a better solution is to lobby github first before doing a cancel campaign
(2) ice contracts: this is bothersome; but I can see how it could be a bit complicated given ownership by MSFT and the complexity of government contracts. It is bothersome though.
(3) closed source: I don't put a lot of weight to this criticism. Just because I can't run Githubs code for myself.. I mostly shrug. Yes, I'd prefer for it to be open source too, but I respect there are various for-profit models out there (and holy-hell I wish I was payed market-rate for FOSS work).
(4) no self-hosting option: Seems like the same point as (3)
(5) CEO leadership discrediting copyleft: bothersome, but without concrete examples, for me it is not fully substantive and just bothersome (but not major, like, wow, they are shutting down FOSS projects, or aiding in getting them taken down by ginning up BS charges, etc..). So, yes, the CEOs of Github were at times discouraging to copyleft. Are they evil incarnate here where those bad actors needs to spurn everything Github? Did the CEOs of Github personally oversee any FOSS projects being sued, or made into non-FOSS? Did they personally increase the cost on FOSS?
(6) owned by MSFT: big companies are big companies and they are really hard to avoid completely... MSFT has had a big culture shift in the last 5, 10 and 15 years. It's not the same company it once was. That is not to say this is not a concern. Though, until I have specifics around how/why I thnk MSFT has become actively evil, this remains just a notable concern.
> If you had actually read the whole article[1]
Apologies for seeming like I only commented on the first paragraph. A lot of the article seems vacuous to me and generally trying to gin up a mountain out of what might just be a gnarly hill.
Git is confusing for just about everyone. It is really easy to shoot yourself in the foot once you step away from git add/git commit/git push. Hell, you can foot-gun yourself even with the usual workflow.
GUIs help with bridging the gap, but because the GUIs make everything that's actually happening opaque, troubleshooting gets really complicated when something goes wrong (GitHub also sells professional services).
Github and Gitlab have done a lot to make Git easier.
SourceForge went bad ... and everyone left. That doesn't seem like a bad thing and there's no reason for me to think any given site / service will or won't go bad too. I expect that for any number of reasons I might need to move from one site to the next.
The rest too is kinda hollow to me. The fact that they're for profit doesn't upset me. I figured they wanted to make a profit when I signed up even ... not sure how that would surprise anyone now.
Co-pilot, I personally don't feel there is a compelling reason to leave github due to that either.
Maybe I'm not versed enough in some of this but as a rando dev I'm just not having any problems on github these days ...
I'm not saying the author is wrong or right or that I'm right or wrong, just that I'm not finding that article very convincing.
Also they kept changing the services they offered to projects so often I gave up on that as well.
And then it killed itself in its confusion.
FOSS projects like the Linux kernel use the GPL license because the developers want their code to be free not just for themselves, but for everyone everywhere for all time. It's not acceptable terms for you to take their work and use it to build an alternate operating system that you aren't going to share. If this wasn't important to them they could have just published their code under MIT/BSD licenses.
If you were to build an AI that used the Linux source code to generate a "new" closed-source operating system, in a very real sense all you've done is invent a new way to plagiarize the Linux community's work so that you can weasel your way out of their license terms. Even if you got away with this in the courts, it's obviously very unethical.
What Copilot does is enable the mass plagiarizing of open source code from everyone all at once, mixed up together so that it's hard to know who the original authors were, and then pretend that somehow this makes it ethical.
Anyway...
I have always been curious as to why the largest hosting for OSS isn't open source itself. Maybe I am not intelligent enough to realize the reasoning behind this. Imagine if git wasn't open source! That's like an OS that only runs open source software but isn't open source itself. It just doesn't make sense for people to trust such an obviously flawed service...and yet they do.
And if that wasn't enough when GitHub got acquired by Microsoft very few people thought it amiss. Indeed, even now a lot of people are happy that Microsoft is running their digital homes. I think if GitHub was measured on the FOSS scale it would fall short on every measure.
Co-pilot isn't even that big of a deal. That's just the icing on the top. GitHub is untrustworthy top-to-bottom even before there was Co-pilot.
But I suppose convenience always trumps openness and freedom. It's especially sad because the whole point behind FOSS was this. Is the whole FOSS idea getting old?
The worst thing is that GitHub has monopolized the open source world. We can't even think of moving off of GitHub because of what we'll lose. But how about we do this:
We create a dummy repo on GitHub for our project that has all the fancy README, releases, issues, actions etc. but we keep the actual code out of GitHub on an open source service. Would that work? Is that feasible?
Basically, we use GitHub's wide adoption for what it's meant to be used: to market/share your project but keep the source code on a separate platform. This would create a new host of problems but I think it can actually work.
It seems so---I read the post above saying essentially "I don't care about any of this criticism, I'm going to keep using GitHub because switching would cause too much churn" and did a double-take.
The whole point is to punish the boycottee which may cause inconvenience. Much easier to dogpile on twitter I suppose.
I can’t really think of a universally sustained boycott since South African apartheid to be honest. Even the current Russian stuff is too inconvenient for the majority of the world so they keep buying their oil and gas.
I use Github to publicly host my projects. I use github as I want them as public as possible and that is where all the people are.. I.E. it has the lowest friction for others. The tradeoff is a bit of extra work on yourself to make sure you keep the 'git' part of your repo the central thing. Use the hosting as extras appropriately, but don't rely on them. All relevant context, reasoning, etc. needs to be in the git commit messages. The git repo should stand alone and tell the complete story.
IMO following this simple rule you can keep your project git repo whereever is easiest for the users without lockin. Github's added features over basic git hosting are decent but none of them are irreplaceable if you keep your repo up properly.
I have my own gitea for private projects, but the open source one I host primarily on Gitlab. It's where I set up the CI, it's where I have pages, docs, etc. I do have a mirror on Github, but on the "Contributing" section from the README I make it clear where I prefer to receive PRs.
What OP was criticizing was these larger FOSS projects who don't seem to mind that they are doing all their work on a closed platform, and that they have a lot to lose if Github decides to pull the rug from under them.
Consider Macbooks...
They seem to be making slow but steady progress on this at Gitea, maybe at other FOSS forges, too.
Great, you already lost half of them.
As for the reason, it's simple: the process of submitting patches sucks. I wrote about it: https://dcz_self.gitlab.io/posts/git-botch-email
This would be like someone trying out make the first time and not realizing why it isn't working becaue they didn't realize they need literal tab characters in the make file for the rules to work. But if they don't read the documentation, there's no way they would know that.
The real problem is people trying to figure out how tools work by experimentation as opposed to reading documentation. If someone reads the documentation of git send-email and the project's contrib document contains the preferred settings for that utility, then submitting patches should not be an issue.
I have limited mental resources, and given the choice between a tool where I have to spend half an hour before I can begin using it, and a tool which will guide me, I'll always choose the latter. After I'm done, I can even forget I ever used the latter tool! It's a boon for one-offs.
Keep reading, there's more criticism on other aspects of the tool.
That would address one of your concerns about saving the email on disk and also ensuring that the headers have the correct contents.
If the project's contrib document contained information about what settings to use for format-patch and send-email, then the process would be much more seamless. I haven't looked at the kernel (or subsystems) documentation on that. The git project itself doesn't seem to contain that information though.
Regarding your other point about using your email client to handle sending the emails, git does have a utility called imap-send that would allow you to upload the patches to an IMAP folder, which, I believe, would allow you to then send the messages using your MUA of choice instead of git send-email.
[1] https://www.kernel.org/doc/html/latest/process/submitting-pa...
If you really want to contribute, but don't want and e-mail based flow, send me a mail, and we can discuss.
It also builds on top of ActivityPub which is supposed to allow federation with the greater ActivityPub ecosystem ("fediverse"). I guess this would allow people to like or comment your issue or pull request from Mastodon and those other platforms.
I am not that excited about that last part, but federating pull requests sounds like a killer feature and a necessary step for a chance to topple GitHub. If hosting my own forge means I have to either get patches over email or let people register so they can create their own fork here, it's a non-starter for many.
edit: here is Gitea's issue: https://github.com/go-gitea/gitea/issues/18240
Git already has a pull request feature [1] that's as federated as it can get. The 'request-pull' command can be used to request pull on upstream repositories hosted anywhere (or not at all). The only requirement is that the downstream clone must be online. I know that HN isn't particularly found of email-based workflows for git. But requesting a pull is as simple as copying the output and mailing it (or via any text messaging service) to the maintainer. And doing a pull as a maintainer is actually easier than doing local PR merges using Github.
I would think it's not just HN. Do you think people are using github PR feature because they just don't know about the email-based workflow available, but would prefer it if they did? Most people don't want an email-based workflow here.
I love email based workflows as much as anyone but for many people this is not "simple". For one, you need to be able to send plain text email or at least not have your client mangle it too much.
You can replicate that with email with a patch-based workflow, if you have a mailing-list server with public archives. That is not that much less software, and you have to deal with email deliverability etc.
Simply sending someone a git-request-pull doesn't carry anything for posterity. It contains a link to some place that hopefully contains the changes at one point (if you typed it right, git does no validation), but probably won't contain them for long.
There's a reason that's the case, and it's likely one of the reasons I should just use github as well; despite being morally opposed to what they're doing WRT copilot.
There's a bug for that: https://github.com/go-gitea/gitea/issues/1029
Sibling has already provided the tracking issue for getting off it; as far as I remember it's close now :-)
Better start self hosting gitea right now. And you can do it for free. I think the best option is oracle cloud, or does someone know how well it works on fly.io? https://paul.totterman.name/posts/free-clouds/
There are great FOSS tools for hosting source code like https://gitea.io/en-us/ and https://codeberg.org/. I’d put the work to self-host and even contribute to any of them.
This makes no sense. Our GitHub lock-in is due to the audience. If a FOSS project leaves, they lose the GitHub audience. It's a social network.
VSCode is absolutely nothing like that. You seem to be condemning Microsoft for making a really good product.
Is JetBrains also evil for their suite of amazing, convenient products?
However - the main gripe this community has with VSCode is that it is only partially open source. Microsoft adds extra bits to the code that is not open - mainly the items around telemetry. You can get around this with VSCodium or VSCode-OSS, but those are arguably forks and not a MS product.
My criticism would only be that the telemetry is not obvious enough for a casual downloader. I don't use VSCode but I'm sure a lot of people do without knowing about it.
It literally just means "has problems."
Then it drops this:
> GitHub's business model has always been “proprietary vendor lock-in”.
How is this Github's business model? Unless you're using Github specific features like Github actions and workflows, it's fairly easy to switch to another Git based host.
Then the article provides "alternatives" that are all lacking important features.
> If you're ready to take on the challenge now and give up GitHub today, we note that CodeBerg and SourceHut0 are excellent options right now.
The article immediately talks about drawbacks with all of these alternatives and then mentions a guide on how to self host using git lab. Why would I go through all the trouble of swapping to a different version control host if I don't gain any value? In addition to not gaining value, I'll also lose features that are very nice to have.
This article doesn't convince me at all. Yes Copilot is questionable and we should pursue the ethics behind what it does, but if you want to convince people to give up Github you should at least be prepared to give an alternative that offers a great deal of feature parity.
> if you want to convince people to give up Github you should at least be prepared to give an alternative that offers a great deal of feature parity.
These can't both be true. In fact Github is hard to switch away from, because of all the Github features. This is the lock-in. Github then monetizes this by charging for large files (https://docs.github.com/en/repositories/working-with-files/m...), private repos, etc. So the SFC argument is that you should switch away from Github now and get the alternatives to feature parity, to avoid Github getting a monopoly.
It’s easy to say “our rights are being stripped away” but the view that businesses should operate like non profits or government services with the common good in mind is ludicrous!
I don’t think that that is what is being asked here. Even if GitHub didn’t violate license terms, it would have ways to make money.
GitHub provides a place for people to easily collaborate on FOSS software. It has helped millions of people getting into software development or into their first FOSS project by lowering the barrier of entry significantly.
Can you imagine how many people would start contributing to FOSS early in their career if they had to deal with mailing lists, patch sending, multiple git remotes, rebasing, etc. all at once just to start providing a small contribution? - Probably not as many.
I don't support everything GitHub does and I do see CodePilot as problematic but the article opens with "Those who forget history often inadvertently repeat it.". - You know what has screwed us over a lot in recent history? Cancelling something/someone without thinking it through first. Oh the irony.
There is a difference between 'how things are now' and 'how things could be'. Imagining and wanting a different status quo is not by itself ludicrous (especially since we all stand to benefit from such businesses), it's a first step towards change.
It's hardly the fundamental principle. The fundamental principle is that people need things to survive and its more efficient if people specialize and trade than if everyone creates everything they need.
The pervase idea that businesses should focus solely on generating profit is also directly responsible for lots of problems almost anywhere in the world from driving out less vicious competitors to rent seeking to externalizing costs to everyone else e.g. via pollution.
Fairly self-evidently, the sane fundamental principle for a business is "make a good/provide a service, and if you do so well, you make a good profit".
Unfortunately, for the past few decades, businesses in the Western world (and particularly the US) have increasingly been operating based on a fundamental principle of "make as much money as you possibly can, and if you have to make a good/provide a service to do so, that's a necessary evil".
But I'm perfectly happy with GitHub and I'm fine if their ML thingy makes money off my code, I get free actions runners, a nice UI, pull-requests, etc, into the bargain, not bad.
Like knock yourself out working out if the "monkey selfie" Supreme Court case law applies to copilot or not, what jurisdictions it covers, etc. But I don't care, sorry, I'm not interested.
As an indivdual you can certainly think so. But as a community we must balance how much advantages do we get from GitHub versus how much advantages it gets from us.
Considering that GitHub will probably make millions with copilot, then it is fair to say that a big part of that success comes from the quality of the code we collectively put one. Therefore, money should be shared. And the first thing to do is to ask Microsoft : how much money do you make on us ? And the only possible way to get an honest answer is to make sure the management of GitHub is done jointly by MSFT and the community.
I don't think this is going to happen anytime soon. So I'm seriously considering getting out.
But you know what? Even after copilot my code is still there, for people to make money from (as both GH and random people already do), to learn from, to cut up and rehash, to reference, to write (much) better versions of, to generally advance humanity. I know this is the root philosophical conflict between free software and open source but I wanted to state my view since this is a call to action and we'll be seen as betraying some ideal for failing to comply.
Isn't this sort of the expectation of these free services? They provide a service to you for free, and in exchange they are able to collect data from you, store it in a database, and do things with it.
I fully expected GitHub to do something like this (minus Enterprise repos that pay big money), and it's why I stopped using them. I instead use gitbucket, which is a free and open source self hosted thing.
I also expect GitLab to do something similar. It is as they say in the crypto world.
Not your wallet, not your crypto.
Not your source control server? Not your code. Legally? Might be, but that doesn't matter. Do you have lawyers?
Otherwise that's the price of open source, you don't get to choose.
Licenses don't need to forbid re-upload generally to be incompatible with GitHub. Given that uploading to GitHub, according to someone quoted on page 1 of this discussion, grants GitHub the right to use that code to "improve their service" (whatever that means -- maybe Copilot?), not explicitly granting that right to them, or the right for someone else to grant it to them, is enough to make them not allow uploading to GitHub.
So that should be more like: "If the code had a license that didn't specifically allow Microsoft to use it to 'improve the GitHub service' you could order GitHub to remove it."
And then slowly they started changing policy, closing the source, adding ads to the platform, .. then one day you realized you've locked yourself into their ecosystem. It's a bunch of work to move away. It will be worse with Github since they bolted so many things on top of Git.
Copilot’s contention is that it’s exempt from copyright restrictions under fair use doctrine, which means that your license that says they can’t use it is irrelevant and legally void.
AFAICS you don't even have to: GitHub reserves the right to use uploaded code to "improve their services" (in some unspecified way). The GPL doesn't grant that right -- i.e. the right to grant Microsoft that right -- to any license holder, so anyone else but the copyright holder will infringe on their copyright by uploading it.
> But I don't care, sorry, I'm not interested.
If you've ever uploaded, or are planning to ever upload, anything containing code written by anyone other than yourself under a FOSS license to GitHub, you probably ought to be interested.
* They make money from Co-Pilot? Great!
* They sell software to ICE? Good. Why wouldn’t they? I’m not interested in anti-immigration-enforcement politics.
* They’re a closed-source for-profit company? Great! That’s why their product is high quality.
OK let me take your projects and make money from them and not give you anything in return, not even credit.
> They sell software to ICE? Good. Why wouldn’t they? I’m not interested in anti-immigration-enforcement politics.
ICE policies aside, you should be interested in immigration politics because they're important to people and businesses. Apathy is a bad thing here.
> They’re a closed-source for-profit company? Great! That’s why their product is high quality.
hahahaha, there are plenty of open-source not-for-profit companies with high quality products and many many many more closed-source for-profit companies with terrible products. See also Microsoft Windows and Microsoft Office.
Burning coal for generating energy was very "successful" and had great adoption. It's still a terrible method for the environment, and we can feel free to regard it as negative.
In this analogy, Microsoft products would be like generating energy by burning garbage, aka a dumpster fire.
> because burning garbage for energy when the alternative is a US style landfill is literally better in every way.
Now that you say it, I do remember reading that Sweden burns garbage for energy. I would have thought that the main problem would be arbitrary emissions from the plant, but from the plan of a typical one, those are trapped and/or filtered [1]. I still think that in the US this would be harder, since people are probably prone to throwing more things in the garbage than they should; and don't recycle as much as the Swedes do.
----------------------------------------
[1] https://www.americanprogress.org/article/energy-from-waste-c...
I am lucky that I can avoid their products. But millions of users and orgs can't because that's what others around them use. So they're stuck with it no matter how much worse it is than open-source alternatives.
(and yes, that is trivially simple. If your software can't do that in a performant way, you don't understand what product you're trying to out-compete)
It's not that there's no alternative, it's that the alternatives are just not good enough when it comes to "I need this stupidly complicated thing, thanks to 20 years of spreadsheet formula history at my organization that I have zero power to even remotely change, done in seconds. Not minutes".
This is one hundred percent okay with me and many other open source contributors. It's a no-strings-attached donation to mankind, and if someone else finds value in it, you don't complain, you cheer. Who needs attribution when people are actually _using_ something you made; you saved someone a good deal of time trying to write a solution themselves, perhaps.
That's fine if your project says that. But there are a very many of projects which specifically say otherwise.
edit: Not implying you're wrong, just moreso that there's a large chunk of people who aren't going to be motivated to fight on your behalf. I understand there are some useful areas such as hardware drivers for GPL, but it's simply not an IP constraint that sounds at all fun to work on as a volunteer.
I guess this just all seems rather activist to me but people aren't seeing the big picture; our jobs as programmers are about to change in a very big way. It won't be long before more competition enters the space (e.g. Salesforce and Amazon) and ultimately it won't matter if this model saw some GPL/MIT code because the next one will work twice as well without having seen any of it.
You will at least save some people some effort they go to in order to collect and provide attributions to every piece of code used.
So? Is "activist" supposed to carry a negative connotation, or what?
Like, maybe you could all get together and learn how machine learning works and train your own clean model? That feels far more positive and in the spirit of progress to me. I guess from my point of view it's clear - no laws will save you from the next wave of large language models. The weights are trivially distributed meaning once it's trained it's more just a fact of life we all have to deal with. So making demands when you have basically zero leverage rather than admitting defeat and working within the new constraints of progress.
I just feel like it's an inherently philosophical position and people are acting like code theft hasn't been common practice for the history of all software.
See git itself.
That's a huge stretch for a few reasons. First, the reliability of GitHub.com is very poor and it has had tens of incidents in the past few months. Second, their hosted solution, GitHub Enterprise, is notoriously poor, difficult to maintain, always late with features. Third, their main competitor, GitLab, which is open core, was kicking their ass for years with features and quality until GitHub got the unlimited funds of Microsoft to be able to even come close. More to the point, there are tons of good quality open source software, and tons of really poor quality closed source. Openness of code matters little for quality (besides the fact that with open source at least you have the option to see why and fix).
Regarding Copilot, they're ignoring licenses, training models on everybody's code, and selling the result as a service. Sounds very sketchy.
- what you opposed negatively affects your target audience in real ways that matter, for their career, livelihood, or some other means
- Show that continuing to be apart of an old model will be damaging in the long term
- Also importantly, the thing you are referring people to do needs to be seamless. For instance, they mention SourceHut, and I sure hope SourceHut has all the core features and ease of use of GitHub, because if not, you are likely already going to lose in this conversation to most
Without factoring these things, its great to point out issues, and rightfully they should, but its not going to mean much in terms of action
yeah these same types of purely philosophical arguments have led me to move away from services like DropBox and Google Drive at my own expense. Lost too many hours trying to find the “we do things the right way” alternative—when it comes down to it, if it works well and keeps the friction low, mf’s can have their bag
Direct competition against GitHub is barely worth contemplating; the practical path I see for replacing GitHub must be more indirect:
1. Some nonprofit foundation spins up a GitHub alternative focused on transparency and strong data privacy protections. This alternative has only a small fraction of GitHub's most crucial features.
2. A major FOSS project---which values the principles of the new alternative over the practical benefits of GitHub---switches away from GitHub.
3. Satellite projects reexamine their use of GitHub and slowly start switching over as well. The new hosting service incrementally adds features in response to demands from the growing userbase.
Steps 1 and especially 2 will require motivation by philosophical arguments, even if I agree that the linked article's execution wasn't perfect.
Simply, I'm highlighting what I believe to be crucial in giving the philosophical argument some teeth in purpose and next steps.
If the friction cost is low to do the right thing, then doing the right thing becomes extremely palatable
My reasons though are not because of Copilot. It is because I do not use git for my own projects (I self-host Fossil and mirror on Chisel). But if someone else makes mirrors of my code on GitHub (or CodeBerg or SourceHut) then I do not have an objection to that (making more mirrors on different services may be better, anyways; unfortunately if you are using git then a header will be prepended to the file before computing the hash, which makes the integrity more difficult (although it is still possible, since the header is predictable (as far as I know))).
Seeing a few examples of output from Copilot (although I have not used it myself and do not intend to do so), they do not seem to be a very good quality. So, I think that it is not worth it, even regardless of licensing issues.
If you do move your project to another service, you should please use one which does not require JavaScripts enabled to be able to view the code (even if other functions do not work). Ideally it should also work without CSS. (For these reasons, GitLab is not acceptable.)
Previously it was free and the noises that GitHub had given made it sound like it wasn't finished. Turns out it was.
You don't know me, but it's totally fine to drink this this mysterious blue liquid. I pinky promise it isn't toilet cleaner.
But it gets worse: because you can't tell whether copilot is giving you illegal-to-use code, even just looking at its output can make you liable for future transgressions because seeing copilot code may expose you to license-encumbered code that you might now use as inspiration for new, non-copilot code. Even if you in good faith believe your new code is your own, the law does not agree with that assessment: if your code is similar to something license encumbered that you read in the past, even if you forget about it, you are now potentially in legal trouble.
Simply by using copilot at all, you're taking on a risk that is great enough to go "actually, I am not even going to try this thing".
Another critical reason is that GitHub is proprietary and has a pile of services with lock-in that (unlike git) aren't easy to move elsewhere. Some people depend on GitHub for their livelihood (through GitHub Sponsors). Some people's software integrates tightly with GitHub bots, or Actions, or issues, or project management. Everything other than the code itself is incredibly difficult to port over to another service.
And no, I have not switched (yet) to a non propietary git system, I am using Bitbucket (yes on a Raspberry PI) since it has plenty of already built in integrations, sadly they ended their offer of self hosted for small teams. It used to cost 10 bucks to get a licence to host your server for 10 users... lord most OSS is a lone effort or surelly under 10 guys, i am sure they would have continued that offer given they had more small customers. Anyhow there are plenty good alternatives as Gitea that someone might use, and that would be the true intent of git as a decentralized platform... nontheless while software is one of the best paid professions nowdays it is full of cheap (greedy) people
I think if Microsoft was figuratively and literally rate limited from accessing such a huge swath of open source code, they might not have been the ones to build something like Copilot and we may have been better off for it. A different, more ethical team might have made something better. Maybe it would even be an open source project.
That's my point-- how? Presumably you're just going to host your FOSS somewhere else that Microsoft can still get it.
Or, the platform could make it part of their terms of use that licenses must be respected even for machine learning. It could be enough of a barrier to dissuade them.
Would something like that stop Microsoft? Maybe, corporations tend to look for low hanging fruit for stuff like this.
Small changes in a situation like this can make a big difference in outcome.
The problem isn't that Microsoft can access you code. It's the section D.4 on Github's terms of service [1]. If you are hosting your code on Github, you are essentially giving Microsoft the legal rights to do certain things with your code. That section is broad enough to allow Microsoft to train copilot with your code. It gives them the freedom to disregard the license. On the other hand, it would be a clear violation of the license if they read your code off another hosts and reproduced it verbatim. I don't know whether any legal challenges would stand in a court, but at least you are not giving away the right to do so.
[1] https://docs.github.com/en/site-policy/github-terms/github-t...
Also, nothing prevents Microsoft from completely ignoring licenses and terms of use elsewhere. This product isn't open source: they can keep the list of projects they crawl to themselves, and I don't see any way to stop it.
What they're doing does feel icky, and they could have mitigated a few concerns by a) making the inventory of the full training set public and b) at least attempt to attribute if there is a direct copy (which by their own admission happens about 0.1% of the time). These seem very simple steps they could take, and takes away the "shady behavior" argument.
IIRC they just block that suggestion if they detect it's a copy now
Single 6mb executable with a version control system, web server, bug database, forums, import/export/sync with git, repo browser, much saner CLI than git, etc. Been using it for 2+ years for all my projects and love it.
Anyway, why would I usee Fossil rather than Git exactly? Are there code hosts that support it? It does look interesting, but I don't see the motivation to do it when so many hosts support git (people who think GitHub is the "only" host need to look around: there's literally dozens of options), and I'm quite ok using it specially with intelliJ and magit (emacs) support.
If you use any open source operating systems a significant portion of software you interact with are C (kernel, shell, utilities, sqlite etc).
1. Hashes are computed without adding any additional headers, so it will be the same as computing the hash normally.
2. The /raw capability is a good thing to have.
3. The deck format is not bad.
4. The command-line interface is less confusing than git.
5. It is written in C.
It isn't perfect, but it is more than good enough. (If I do have to change it, I would do "Generalized Fossil" (I already wrote the specification, although no implementation exists yet as far as I know), which will have the same five advantages listed above, and is compatible with the same /raw capability and deck format of existing Fossil repository, too (except old technote edits, but fortunately I do not have any).)
See [0] for the Fossil deck format. (Generalized Fossil uses the same format, but any combination of cards is allowed, as long as there is exactly one Z card, and not more than one W card; they must still be in the correct order with no duplicates. There are many other details too (e.g. "subrepositories", which can allow you to optionally make decks unparseable in some subrepositories), most of which will not be mentioned in this comment, but I will say that it is mostly a superset of the ordinary Fossil format, except that ordinary Fossil allows cards to be in the wrong order in some circumstances (specifically, technote edits in some versions) that Generalized Fossil does not allow.)
See [1] for an example of the /raw interface (in this case, a mirror of one of my own projects; however, this will work on any publicly accessible Fossil repository). The name "trunk" at the end of the URL is the branch name; you can compute the SHA-1 hash of the returned file and substitute that in place of the word "trunk", and you will get a permanent link to that version. The lines starting with F are the files in that version; each line has the file name, and then the hash, and sometimes another field specifying file mode ("x" means executable). Substituting the hash of the file in the end of the URL will access the contents of that file. The line starting with P has the hash of the previous version; you can put that in the URL to access the previous version.
(In at least one case, I have used this /raw capability to download a single version of a Fossil repository. It is a simpler interface than using /xfer, if you do not have Fossil installed on your computer.)
[0] https://fossil-scm.org/home/doc/trunk/www/fileformat.wiki
[1] http://chiselapp.com/user/zzo38/repository/freeheromesh/raw/...
Just yesterday I had to explain the basic premise/history of Git to a young intern. I had asked him if he was using Git to manage his little pet project the company gave him to play with. “No”, he replied, he didn’t know what the company’s policy was to posting code in public on GitHub. As I explained to him that “git init” was all he needed, no GitHub or even no repository on our local GitLab was necessary, his eyes grew wide: “But how does that work??”
I’ve had to explain this same thing to multiple novice devs of various ages. It baffles me. I consider it one of the greatest ironies of software development today.
It’s like explaining to people that they could just talk to each other using a thousand means, instead of having to communicate by netcasting at each other through some shared social media platform.
https://git-scm.com/book/en/v2/Git-on-the-Server-The-Protoco...
".. you don't have to keep the previous versions, so it saves space? OK... No, I was talking about the real git."
My experience with younger folks that have never been outside the Windows world has been that it was a lot harder to make them understand. Young folks with a unix-y background were much easier.
1. Base technology arrives and is adopted en masse
2. Some entity wraps it in an easier-to-use interface
3. After some time, that interface becomes the de facto standard
4. Developers who started developing after #2 and especially #3 don't understand how the underlying technology works or even that the wrapper is just a wrapper
I'm reminded of how many Juniors I trained that didn't know jQuery was itself just a Javascript library and not a language in and of itself. None of them knew much of any of the underlying Javascript it was wrapping. I'm seeing the same scenario play out with React and Vue right now as well.
Yes, sure, that could be one useful strategy.
But no, you can't demand it as some kind of mandatory component for a project to "count as a real FOSS project".
And for git, especially... It's not as if an easy-to-use interface had ever been a priority there. ;-) What one could ask for from them, though, IMO, would be a website of their own where they at least explain what it is. Because if such a thing exists, I must have missed it. I mean, git-scm.whatever isn't that, is it? That feels more like... Idunno, it's actually Atlassian in a thin disguise, isn't it?
Since then I've been obsessed with knowing how things work under the hood and doing as much as possible by myself until I understand what's being abstracted away from me.
Git wasn't even invented yet when I was in school and I did not learn about SVN until my first programming job in 2001 where a senior explained it to me.
Most people will get that.
That was a bit shocking to me. I started with local VCS only all the way up through school and work.
Once, I had to introduce most of the employees of a company to git and git hosting. They were trained from scratch to use git reasonably on the first day. Use of code hosts were taught only on the second day. They were taught hands-on on our internal host and on public server. None of them had any confusion over differences among git and different git hosts. While they did make some mistakes afterwards, none of those were based on total misconceptions.
If you didn't already know git was independent of the SaaS products, you'd have no reason to suspect it was independent.
"Well of course, MS invented Git!"
Wrong on so many levels:
1. Conflating Github with Git.
2. They bought it, not invented/founded it.
3. Most importantly, MS's recommendation to use Git for TF existed long before they bought Github.
Needless to say, I didn't join that team. Unfortunately, misconceptions like these, and a refusal to use Git[1] due to its complexity are quite common at the company - and it is one of the larger SW companies[2] in the country.
[1] I'm OK with any DVCS. Even I prefer Mercurial to Git. But most of the company prefers SVN or TF's version control. In 2015, many teams had to be dragged kicking and screaming by IT from CVS to SVN. In 2015, when Git already dominated the world.
[2] By number of employees with "SW <something>" in their title. Not by revenue, etc.
I kind of like this friction. If a software developer finds git too complex, that’s an important signal that I should minimize my exposure to them. I’d never work in a team that dumb, and I wouldn’t want any dependencies on their code.
Of course there are other reasons not to use git, but complexity is not one.
Not git's fault, though.
The main issue is repository size, which is hard to get from bad workflows alone, with the exception of insistence on tracking large binary files in Git (rather than LFS, DVC, etc).
Doesn't the uptake in Github which centralizes this distributed system kind of invalidate its main tenant?
I haven't been a hundred percent sure the overhead was ever worth it at most other types of paid gigs over the years. Adding complexity without value is a mistake imo. Maybe it's my own fault I haven't seen the value...
Indeed, even when forced to use SVN, I would simply check out the SVN trunk, create a Mercurial repository in that "working" directory, and then clone from their whenever I did any development (one clone per feature). I would then push back to the main Mercurial repository, and push that to the SVN server.
Other than that - true. No real DVCS benefit compared to SVN. I would imagine most of the nicer Git features that people use have, or could have, analogs in SVN.
Once did something similar with Git, the one time I was at a place that used SVN. I didn't trust any of the git-to-SVN tools so I just did all my work in a local git repo, then copied my working directory to an SVN-controlled directory and committed maybe once or twice a day, when I had something worth preserving.
The ~week before I switched to this workflow was nerve-wracking. Not being able to make all the branches I want for any purpose at all without it showing up for anyone else, or to make shitty commit messages for my own local junk-commits before it was ready for consumption by anyone else, was awful. Having to worry that every little vcs operation might mess up someone else's stuff was the worst.
Luckily I was working on an isolated part of a larger system on my own, so this worked OK. No incoming code to worry about, so SVN was write-only from my perspective.
Yes, unless you are working on a tiny number of source files. The value is that you can work on entire copies of a code base at a time instead of single files. Git does a much better job than making lots of local copies of source code, passing around big diff files and is a lot less hassle than older centralized source control systems.
> Doesn't the uptake in Github which centralizes this distributed system kind of invalidate its main tenant?
No. Not at all. Github is actually a peer with an a few integrated extras for managing tickets and requesting your code be merged into github's repo branches. Most of the centralization is around access control and automation (i.e. continuous delivery, unit tests, etc...).
> Adding complexity without value is a mistake imo.
The main value of GitHub is reducing some operational complexity. For a small 1-4 person team, it may be of little value, but for larger teams, access control, issues, and automation can have a lot of value. For really small teams, fossil is actually quite nice. That said, there are a lot of great alternatives to GitHub that give you similar features.
> Maybe it's my own fault I haven't seen the value...
Probably not. I did the first 20 years of my career without source control and was able to build some pretty big applications. I do think dvcs was a big improvement, and I'm glad it exists now.
It’s useful because it makes it easy to have multiple copies of a repo and merges and branches are trivial. Just for my own project I might have clones on multiple machines. When I mess up and make a change, it’s not a big problem and merging is easy.
For team collaboration, I find it better than centralized because it makes it easier to work on multiple branches. It also encourages merge requests from people outside the team.
I'm pretty sure half of the teams that find Git too complex wouldn't find Mercurial to be complex. But they haven't heard of Mercurial.
However my own projects always start with hg init. In addition to the UI being less insane it also feels consistent and nice to use.
At one point I was on a team using Darcs, which was another level of beautiful but clunky in many ways. It is why I'm rooting for the Pijul team. [1]
Given the context of this thread, comments like these are part of the reasons teams don't use Git.
One probably needs 10x such teaching moments with Git compared to other VCs - including good DVCS.
And I also take issue with calling its “non-100% ubiquity” a problem. There is plenty of space for other tools. If anything, there’s almost a “git monopoly”.
Sure it could improve in parts, but the general discussion here doesn’t seem to be focused on improvements at all…
The practical approach was to write a simple GUI which supports the operations "Pull" "Stage" "Commit" "Push", which makes live quite easy. And otherwise, try to keep out of anything more complex, until really neede. Works quite well, but I cannot say I am entirely happy. There are some benefits though, due to our company using a GitHub Enterprise installation for all version control. Ironically, our group recommended that some years ago, but because GitHub comes with a lot of neat features, less so because our love of git :)
And yes, over time, I am picking up more of it, but I try to do that on a slow pace, because I consider VC a tool which is important for work, but which shouldn't take time away from work. And Git definitely takes a lot of time to learn, there are far too many mechanics exposed to the end user.
Also, without knowing which issues you're having with git it's impossible to know if it's lack of baseline knowledge or if you're running into real complicated problems.
If you have a GUI with "Pull" "Stage" "Commit" "Push" buttons, I strongly suspect you're not in the complicated end.
Worst case scenario is that I lose a check in, but even that is rare as I can just copy over, start with a fresh clone, and reapply. I don’t think there’s a significant danger of messing up a project. And again, this is just basic developer capabilities and developers should be familiar with a vcs enough.
I blame developers and even most users for being too scared to use it. I think it’s perfectly normal to be scared of it while using it.
I think all users of technology should have some basic competence. Like every human should be able to be a basic Unix user, every human is capable of using git. And every developer should be capable.
Every one should be able to use Git. However, Git really isn't the best tool for everyone - or even most developers.
That was the GP's point. He acknowledged it's complex. But it's a very clear and very large amount of value left on the table, so people that gets scared away by it will probably practice other kinds of harmful behavior.
(Of course, the option of just using a simpler VCS doesn't tell anything bad about people. Why did we standardize on git again?)
My issue wasn't with their refusal to use Git per se. It was their refusal to find something better than SVN.
Given the choice, I too would pick something other than Git. But still better than SVN.
To clarify, I’m not criticizing not knowing git. Lots of people don’t know git. I’m criticizing the decision not to learn git and being unwilling to learn it due to its “complexity.”
> ... that I should minimize my exposure to them. I’d never work in a team that dumb ...
That sort of arrogance is an important signal to me that I should limit my exposure to the people displaying it. I personally see this sort of attitude as a sign that there may be a dangerous lack of empathy on their side, and I've seen that go south too often.
I find Git's CLI and workflow model complex because I had the "misfortune" of using other DVCS products (Mercurial, Bitkeeper, Bazaar NG, Darcs, etc) before it. All did a far superior job of presenting roughly the same conceptual model in their command-line tools. Git's is a step backwards.
Git didn't invent distributed version control, it didn't perfect it, it wasn't particularly superior to the others [it does have some benefits in terms of performance tho, yes], it was merely at the right place (the Linux kernel when Linus got sick of Bitkeeper's business model) at the right time (when people finally got sick of CVS and SVN garbage.)
And the rise of GitHub was certainly part of the rise of Git's prominence.
Slightly different circumstances and it could have been any of the other open source distributed revision control systems instead.
Different strokes for different folks and not everyone is capable of working together. Comically, people like you would find your comment also arrogant and stay away from both of us.
For what it’s worth, I work on empathy quite a bit and my filter is based on the idea that it saves both me and the other party pain. I don’t want to work in a team where developers are tolerated giving up on super basic technologies like git (what else isn’t tolerated “oh, tcp/ip is just too complex for developers, I give up”) and they probably don’t want me to work with them.
I think one of the best things about technology is that people have great capability to solve problems. Giving up is a bad characteristic. Asking for help is important. Having teams of various skill sets is important. But giving up on basic things instead of getting help and figuring it out is bad as a permanent state for a team.
You'd think developers could at least make the stuff they have to use decent. But no.
porn --> pornhub
best explanation of github for noobs
Though it's frustrating that it's pornhub specifically, because referencing pornography can be awkward or impractical in workplace contexts.
fried chicken --> KFC
"Porn" is a genre of video and YouTube is better known than pornhub.
I'm surprised that at a coding bootcamp students would already have some idea of git ... at all.
New devs - especially those coming from bootcamps (I say this without judgement) - mostly start with practical skills. Industry-standard ways to just get things done. That's how you get a job, that's how you get off the ground. This goes beyond source-control; languages/frameworks, tooling, etc. You enter the territory - with your finite bandwidth for learning - where it's most immediately useful. And then over the years you move out from there, incorporating more and more nuance and detail and auxiliary knowledge.
There's no need for moral panic. "Where it's most useful to start" has shifted, sure. But that's natural; I don't think it's a new phenomenon or in a fundamentally worse place than before. GitHub is a higher-level tool that makes you dramatically more productive than raw git on its own. The details will be filled in as they work their first job.
I wouldn't say that it's "especially those coming from bootcamps". Those coming from uni are no better IME.
P.S. I've seen the professor grading the printed out programs and he'd do it by flipping to the last page which was supposed to have the result output, and then fold the corner of the papers so it looked like he read the whole thing and then put a grade on the top. It was pretty funny.
Not multiple choice tests. We had to write a program by hand and hope to god it would compile without any errors.
Horrible way to teach C imho.
I can't think of a worse idea than subconsciously letting the idea "By default, the way we are supposed store and write code is by putting it in the hands of a deeply centralized third party that will exploit you and owes you nothing" just sort of be the default deal.
There are lots of worse ideas. Things get abstracted from us over time. 99% of the code in active development today probably lives in a source control repo that's in the cloud.
At one point, decades ago, I looked with amusement on devs that couldn't do C/C++. But the reality was that it wasn't really needed anymore for most tasks.
In other words, yes, Stallman was right.
Perhaps a bit of a stretch?
Does the fact that new devs don't know Git fully necessarily mean the quality of code will decrease? I mean, compared with the new devs knowing Git fully...
There's some conflating of correlation and causation here, of course.
That's not the problem. The problem is that they don't know what git IS; they think that GitHub is git.
Not to mention that git is plagued with, as a general rule, utterly confusing defaults for a newbie (and for me, when setting up a machine without my config files!)
I didn't include github, gitlab, or anything else, because we don't use it. The auditor was going off on a tirade about how lack of version control is not okay at all, so convinced they were that 'no github or gitlab' must therefore mean 'no version control'.
The mind boggles. He barely believed me when I showed how git just syncs with other git repos and that's really the start and end of it.
This has actually gotten me into thinking about a few things. What a web site 'backed' by your git repo seems to get you is:
* Some insights to those who don't have a full git dump. Mostly irrelevant.
* CI stuff and hook processing, but this does not need to be done by the system that hosts git, or even a dedicated system in the first place.
* An issue tracker that nicely links together and that auto-updates when you commit with messages like 'fixes #1234'.
* Code signoff/review coordination.
And all of that should be possible __with git__, no?
If you have a policy that all code must be signed off otherwise it isn't allowed to be in the commit tree of your `main`, `deploy` or whatever you prefer to call it branch, then why not just say that a reviewer makes a commit that has no changes (git allows this with the right switches), _JUST_ a commit message that includes 'I vouch for this', signed by the reviewer? And that _IS_ the review?
What if issue tickets are text files that show up in git, to close a ticket you make a commit that deletes it. Or even: Not text files at all, but branches where the commit messages forms the conversation about the issue, and the changes in the commits are what you're doing to address it (write a test case that reproduces the issue, then fix it, for example), and you close a ticket by removing the branch from that git repo that everybody uses as origin?
Then all you really need is some lightweight read only web frontend so that the non-technically-savvy folks can observe progress on tickets in a nice web thingie perhaps, if that. But it's just a stateless web frontend that reads git commit trees and turns them into pretty HTML, really.
Commit hooks to ensure policies such as 'at least 2 sign-off reviews needed before the CI server is supposed to deploy it to production'.
Does something like that exist?
Git has two levels of built-in/native commit signing support. There's "Signed-Off-By" which adds a note to the bottom of a commit. Some projects use that for CLA verification. There's also GPG signing which signs a commit hash.
If you want to use that as a level of "merge request"/"pull request" reviews there's a natural commit to sign that says you reviewed an entire branch: the merge commit itself. You can make a policy of --no-ff merges in your main and important branches. You can make a policy that they are signed (using one or both of the sign off types). You can make a policy that they are signed by someone who wasn't the author of most of the branch's commits.
> What if issue tickets are text files that show up in git, to close a ticket you make a commit that deletes it. Or even: Not text files at all, but branches where the commit messages forms the conversation about the issue, and the changes in the commits are what you're doing to address it (write a test case that reproduces the issue, then fix it, for example), and you close a ticket by removing the branch from that git repo that everybody uses as origin?
There's multiple cool approaches to this that people have tried. Search for "git distributed issue tracker" and you should find some of them. Some have okay web views. There's multiple options for storing the issues. Some use YAML files inside of the branch. The neat thing about files in the branch is that you can find things about where fixes happened using basic branch diffs. Some use git "Notes" which are indeed like git commits as a first class top-level object in git's object tree. Those do have the benefit that they form their own branches outside of your code branches.
It's neat to explore what people have already tried in that area.
A good example is the f/oss community confusing the terms for the Signal.org API and the GPL-published Signal client software.
It's usually not a problem since they tend to already know javascript and can get things working by referencing the node api docs, but it's still really funny to me every time it happens. There's lots of stuff people don't know until they know.
The idea that copilot can somehow “AI-wash” the copyright of large/nontrivial pieces of code seems completely crazy.
Copyright only applies to expression (code) not ideas. It would be extremely hard for anyone to re-implement anything in a matter exact enough to be plagiarism.
For a clean room impl the fear would be patents, not copyright.
Any software lacking patents can (my armchair lawyer guess) safely be re-implemented. If people reimplementing have had past access to the source code of the original, that's probably not even a great danger, so long as nothing is copied verbatim. Ideas/designs/architecture/functionality is not protected by copyright.
I feel like the venn diagram of people who complain about JS being required and of people who have never had to code up a web /app/ that users expect rich interactions without a page-reload is just a circle.
If I'm just reading static text on a web page, theres'a absolutely no reason why I should need javascript to just read it.
100% true if the site is privately funded. In most other cases JS is required for ad integration and analytics.
I don't like it, but I understand that funding is required and ads are the simplest way to get there without getting into the whole micro-payment and paid subscription mess.
True for power users on PCs, but keep in mind that many users use smartphones [0] and tablets nowadays to access websites. The possibilities to block analytics and ads are severely limited on these devices.
JS is also sometimes used to "protect" content from scraping by bots (I cannot comment on how effective this is is, but I've seen it a lot). Again, I agree that JS shouldn't be used like this, but sadly it is.
[0] https://www.statista.com/statistics/277125/share-of-website-...
edit: Without code duplication of course
I specifically called out "web apps" in my first comment as I do understand the value SSR brings to things like blogs, news, or other simple sites where JS is not needed, or where it can have a clean fallback. On the other hand, I write "apps" (sometimes deployed on phones via Quasar/Capacitor as well as on the web) and those get much more complicated. I'm not quite sure how modals, WYSIWYG, rich date pickers, etc translate for a no-js user. Simple navigation is easy enough to grasp but my understanding is that things like NextJS/NuxtJS are really just for first render/paint and then React/Vue take it from there. I could be behind the times on what's possible without JS and using SSR through. I just know the PHP codebase I also work in uses plenty of JS to be functional (not above and beyond, literally "table stakes" stuff).
I'd like to note that Remix does handle everything being tied to a single logic as far as my testing went, I love it. The idea is that basically all interaction is done with html forms (like in the old days) and Remix loads a React bundle that makes that run client side after the page has loaded. It's a very simple model that should work for most use cases, although I don't think it's suitable if you're truly developing a web _app_.
As with everything, balance is key. JavaScript is useful and more appropriate is some situations, and it's not in others. I do hope to see more progress with seamless SSR for SPAs though, I think it would make the internet a much better place.
I have also made and worked on multiple web apps where rich reload-free interaction is expected. In some cases, it has not been practical to support JavaScript-free operation at all, but in almost all cases where JavaScript-free operation has been feasible, I have provided at the very least partially-degraded operation—certainly on all green-field development.
A lot of the places where GitLab requires JavaScript are quite unnecessary, and should probably not have been done client-side at all in the first place, though I’d settle for server-side rendering with rehydration.
But they manage to change and sometimes break the UI in every update and it just gets more and more bloated every day. - At least that's what it feels like.
I'm part of a small FOSS project with about 20 contributors but the changes mean we can now only have 5 max in the Project without paying for licenses or moving the Project to its own namespace and going through an application process for GitLab Ultimate for Open Source (or whatever it's called) which needs to be resubmitted yearly.
I fully understand they are not required to provide services for free but this follows the CI runner allowance reductions - fairly - recently (which is understandable, compute costs money) doesn't give me much confidence for hosting smaller to medium FOSS projects without having to jump further hoops while shouting about how open source they are whittling down their offerings to the bone.
To be fair, there is a valid point here. If even one party has already made their conclusions and enters into a discussion with no willingness to even entertain ideas, instead just to fight their corner, then why would other parties willingly take part? We've all had those engineering discussions where no matter what is said, there are still engineers who refuse to entertain a concept. They're difficult and draining. I can see why the request would be refused if this was the case.
The ask here isn't "don't ever use AI code assistance tools", the ask here is "don't ship something as an AI code assistance product that fails to provide any means of tracking provenance and handling license compliance".
Quoting the post:
> Meanwhile, the work of our committee continues to carefully study the general question of AI-assisted software development tools. One recent preliminary finding was that AI-assisted software development tools can be constructed in a way that by-default respects FOSS licenses. We will continue to support the committee as they explore that idea further, and, with their help, we are actively monitoring this novel area of research. While Microsoft's GitHub was the first mover in this area, by way of comparison, early reports suggest that Amazon's new CodeWhisperer system (also launched last week) seeks to provide proper attribution and licensing information for code suggestions.
Didn't SourceForge make its platform F/OSS again 11 years ago, under the name Allura? https://sourceforge.net/blog/new-projects-welcome-to-allura/ Looks like it ended up at the ASF: https://allura.apache.org/
I feel like that would be a decent compromise for folks not down with their code being used in copilot.
I think underlying discussion should be about licensing, not about website to which you are pushing open source code to. Because that can be easily worked around.
[1] https://docs.github.com/en/site-policy/github-terms/github-t...
most likely they're relying on fair use, which would apply regardless of where it's hosted
GitHub is taking my code and ignoring the license. I don’t understand why anyone would think that is ok.
I find the only people making these OSS claims haven't used copilot and tend to lack any real contributions to OSS. What you're describing is just simply not the case for 99.9 percent of the code snippets being produced/generated based on data from GitHub.
I actually care more about putting code into peoples hands versus someone copying a license file, that's probably why I use the unlicense... "Because you have more important things to do than enriching lawyers or imposing petty restrictions on users"
https://sachachua.com/blog/2015/12/2015-12-10-emacs-chat-joh...
Not.
I mean, you don't have to use these tools yourself, particularly since you have to pay for it and therefore there is no "it is free" allure.
While I understand what you are saying about the pleasure of figuring things out yourself (and do enjoy that feeling myself), I don't feel that the mere existence of the tool affects me that much.
You need to pay more attention to GitHub Sponsors. Not sure how you could write this article and then not even mention it. You're asking people to walk away from a giant pile of money without even addressing it?
> If it is, as you claim, permissible to train the model (and allow users to generate code based on that model) on any code whatsoever and not be bound by any licensing terms, why did you choose to only train Copilot's model on FOSS? For example, why are your Microsoft Windows and Office codebases not in your training set?
I'm not sure I buy this argument. In the first point, the authors state that Copilot was trained on public data. In the very next point, they slightly tweak it by saying training was done on "any" code which loses the distinction between public and private code. Obviously Windows and Office are not public code.
I also interpreted "public data" to mean they trained on codebases that explicitly specified, say, MIT licenses or other permissible licenses. That seems like fair use to me. Those licenses don't explicitly restrict training AI models on their codebases do they? It's ironic if these licenses started banning AI training now though. That would effectively mean Copilot would be sole trained AI model.
I'm happy to be proven wrong though. In general I have a distrust of Copilot. I fear it would make individuals worse programmers in the end at the cost of productivity
So even if it is legal to create a commercial product which outputs GPL code as its main value-add, it still seems like it could put the user in an awkward position of auto-completing big chunks of GPL licensed code into their project.
How do you get the bulk of the users, for whom convenience and features seem to be the primary motivators, to quit?
What prevents Github Copilot from expanding to FOSS that is not hosted by them in the future? Just how Google indexes and caches the whole internet, what prevents this thing from going to public gitea of FOSS projects and scraping to train their model!?
I love Copilot, will pay for it.
I get so many "our terms of service changed" emails and they link to a 30+ page document with not even a diff of what changed. I vaguely remember GitHub sending one out maybe in December 2019 but it linked directly to https://docs.github.com/en/site-policy/github-terms/github-t... and didn't even hint what was different so the only way you'd be able to know what changed is by re-reading every single word.
This is one of those things where you technically agree by continuing to use their service but no one can realistically be expected to read a 30+ page document for the 20 services they do every time a provider updates their terms without a diff. You'd be reading one of these at least once a week.
The email I got from GitHub also didn't include "pilot" anywhere in the body of the email and neither do their current terms of service, so now you need to be able to decipher whatever wording they use to translate back to "co-pilot". After all that I also have lots of emails from noreply@github.com and searching my inbox for emails from that with "terms" in the subject doesn't show anything related to co-pilot.
I'm not a lawyer but I can't imagine if you agree to the terms today but 3 months from now new terms have been added -- you don't passively start accepting those terms without an explicit action to say you do after you've been notified of the changes.
We're self-hosting Gitea and run CI with Drone.
My main concern was exactly Copilot. The concern is mostly principal.
I don't hate Microsoft like I did when I was younger. I think VSCode is a great editor. I still think GitHub is the best social network I know. But I've quit every other social network.
But when your values don't align with an organisation's, you will eventually run into conflicts of interest.
I still push changes to open-source projects that I host on GitHub, but I don't create new ones. The latest project I started runs in a local git directory that will get pushed to a self-hosted Gitea.
When necessary, I'm seriously considering contributing to Gitea to make this transition easier for others.
Gitea is an ~100MB self-hosted binary that mimics GitHub, runs its own SSH server, and it looks GREAT!
There is currently an open investigation case at the Better Business Bureau (filed by my self), but OpenAi refuses to participate. I.E. OpenAI refuses to defend its business practices in front of a government entity.
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
if you move everything to gitlab/self-hosted there's nothing stopping them spending 5 minutes querying the bing index and feeding it every repo they find
You're only better off using GitHub today in the same sense that a turkey being fattened up for Thanksgiving is better off today than one living in the wild.
I think I’ve seen 3 articles about this in just the past week…
> 1. What case law, if any, did you rely on in Microsoft & GitHub's public claim, stated by GitHub's (then) CEO, that: “(1) training ML systems on public data is fair use, (2) the output belongs to the operator, just like with a compiler”? In the interest of transparency and respect to the FOSS community, please also provide the community with your full legal analysis on why you believe that these statements are true.
I mean, I'd be floored if any corporate lawyer let anyone at [large company] answer this kind of question outside of an actual lawsuit. They are essentially asking the opposing team's lawyers to do all this work for them, for free. This is followed by an "obvious[ly]" correct (I'm being ironic) interpretation of the refusal to answer: that MS is wrong but just won't admit it. But go back and re-read the question. The question was architected to produce this impression if it wasn't answered. That's a sign of a bad faith question, rather than a question with intent to learn the answer.
> 2. If it is, as you claim, permissible to train the model (and allow users to generate code based on that model) on any code whatsoever and not be bound by any licensing terms, why did you choose to only train Copilot's model on FOSS? For example, why are your Microsoft Windows and Office codebases not in your training set?
Other commenters have discussed this one already. There is a perfectly reasonable and legitimate explanation here: The do not want to do anything that remotely risks exposing trade secrets, and that's a separate concern from potentially accidentally violating a license. Suppose the model was trained on all these public repos + MS's private repos. Someone else can come along and train their own model on the public code; now they have two code generators whose outputs can be compared to reveal secret information about MS's training set. This time, the article guesses well at the answer: MS cares more about itself than others. Sure. Why would it be expected not to?
> 3. Can you provide a list of licenses, including names of copyright holders and/or names of Git repositories, that were in the training set used for Copilot? If not, why are you withholding this information from the community?
I think this question is bad faith too. It starts by asking "can you". Then, if the answer is "no, we can't", reinterprets the answer as "no, we won't" ("withholding" is an intentional act). It is disingenuous to imply that someone who cannot do something is, therefore, intentionally refusing to do so. In the analysis of the lack of response, the article (finally) admits that it is speculating wildly, backpaddles on the implied claim that MS is refusing to provide this information, and instead takes a different approach: MS scientists can't answer because they are not good scientists. But wait, here's the kicker:
> ... so they don't actually know the answer to whose copyrights they infringed and when and how.
Busted! The authors have essentially demonstrated the question is in bad faith by suggesting that the answer to the question, "Whose data did you use?", is the same as the answer to the question, "Whose copyright did you violate?", which is a logical connection made possible only by the underlying presupposition that MS is totally incorrect in its assertion about fair use in question 1. The framing of all these questions suggests to me that the authors were already firmly convinced of their guesses as to the answers/non-answers _at the time of posing the questions_.
If they actually waited for a whole year expecting a response, that's on them. I'm with MS on the decision not to engage here, even if I share all these qualms about Copilot.
1: What other kind of faith in Microsoft would be even remotely warranted?
2) "perfectly reasonable and legitimate explanation ... risks exposing trade secrets ... a separate concern from potentially accidentally violating a license"
2 a: Sure, they may be separate concerns, but Microsoft is acting -- and you, by arguing for them, at least im- but AFAICS explicitly endorsing their viewpoint -- as if their interest obviously overrides everyone else's. Why should it? They're the ones who want to do this, so why shouldn't their code be the one exposed to any risk? If you want to test if some newfangled house-building material is really as fire resistant as its manufacturers claim, you can set fire to your own house, not your neighbour's. Also, there's only one Microsoft whose interests would be put at risk if they use their own code; but they chose to expose how many others?
2 b: For someone complaining about "bad faith" on the part of others, "potentially accidentally" is some mighty fine weasel wording. What's "accidental" about intentionally building a product and intentionally training it on a bunch of code written by others? (They didn't just randomly press some keys and say "Oops, let's see if it gets trained on our code now, or everybody else's", did they?)
3) "starts by asking 'can you". Then, if the answer is 'no, we can't', reinterprets the answer as 'no, we won't' ('withholding' is an intentional act"
3 a: The word"can" has several valid usages in English. If I say "Can you pass me the salt, please?" and you don't, then you are (assuming you have no severe physical handicap that's stopping you) intentionally withholding the salt from me.
3 b: Even if Microsoft is actually unable to provide the asked for data, the question arises: How come they are? They've built this product. Not building that traceability into it was their choice. Why did they choose not to?
If anyone is "busted" here, it seems to me that's you: Busted as a Microsoft shill.
What case law, if any, did you rely on in Microsoft & GitHub's public claim, stated by GitHub's (then) CEO?
Doesn't Google crawl the entire web? How is crawling code different? Copilot is essentially a more intelligent search engine. The only difference is that people want their websites to be searchable and rise to the top of Google they benefit from the increased traffic. Legally I don't see the difference. As Copilot gets more developed, it may become desirable to have your API at the top of Copilot just like search, and this can help drive traffic to FOSS projects. After all, why would a creator not want their FOSS project to be easier to find/use?
why did you choose to only train Copilot's model on FOSS and not on Windows/Office?
Code from Azure and tons of other Microsoft projects is in the training set. Windows and Office are not FOSS and not on GitHub. Obviously it would be a huge security risk to train on OS code.
Can you provide a list of licenses, including names of copyright holders and/or names of Git repositories, that were in the training set used for Copilot?
Again unfair; this is a gotcha type question. There's all kind of code on GitHub with so many different types of licenses. there's bound to be some gray areas and code that inadvertently made it's way into the model. No lawyer would ever expose themselves to that kind of liability.
The bigger issues I see here are copilot is not free but uses free software and that understandably make FOSS community uncomfortable. However, the Copilot models are incredibly expensive to run interact and Microsoft has to cover the bill. Would it be unethical if Google charged you a monthly fee? Arguable not, because it does already in the form of ads. What Microsoft needs to do is have an opt-out standard like the robots.txt file or noindex meta tag. The problem with GitHub is, unlike websites on the web, not everyone uses Github public repos with the express purpose of having them be easily accessible to the public. Another issues is attribution of snippets is a nightmare, but one could argue that devs do with stackoverflow all the time.