GitHub forking has one big flaw (2011)
zbowling.github.io
zbowling.github.io
You don't want to create an environment in which forkers can easily steal credit from the original author(s). It takes a lot of passion and goodwill for someone to start a new open source project. I think they deserve some credit.
If an 'owner' no longer feels up to the task of managing their project, GitHub lets them transfer it to someone else. That's what happened with ExpressJS and it worked out fine.
Would Xorg still be a fork of XFree86? Would Ubuntu still be a fork of Debian?
Also, sometimes a project is abandoned and a fork is still maintained. Don't you think that the fork should be in the spotlight in that scenario?
OP brings up the example of a project where the original repo shouldn't be the most promoted one, because it's been abandoned, while plenty of other repos are still alive. I blogged about the same thing earlier this month and used the same example.
if you haven't encountered this problem, you will. it is absolutely a problem every developer will encounter. it's only a matter of time.
somebody starts a great project, doesn't have time to keep it alive, and the community fractures because GitHub has no way to differentiate between "original repo" and "canonical repo."
not going to rant about this because I already did in a blog post: https://www.pandastrike.com/posts/20150610-thought-experimen...
If an 'owner' no longer feels up to the task of managing their project, GitHub lets them transfer it to someone else. That's what happened with ExpressJS and it worked out fine.
this is a ludicrous statement. the Express.js transfer of ownership was a ridiculous fiasco full of angry drama, hurt feelings, and core developers resigning from the project.
documented here: http://gilesbowkett.blogspot.com/2014/07/the-bizarre-bazaar-...
also, the idea that you can solve this problem by having the original owner transfer ownership doesn't make any sense. the whole problem is that the original owner isn't paying attention at all, doesn't care in the first place, and wouldn't know who to transfer ownership to, if they did care.
it happens all the time.
True, I should rephrase; from the consumer's point of view it turned out fine :p The project is still healthy.
About ExpressJS, I did read something about one of the main developers not even being aware that the transfer was happening until the last minute. I also heard that there might have been money involved and it probably wasn't a fair process. There are a lot of ethical dilemmas there. A transfer of ownership doesn't have to be this nasty though.
In most licenses I've read on Github, even the very permissive ones, the creator will still have copyright of all forks, even if all his/her code has been replaced.
The problem is compounded by the fact that doing anything related to the ancestor, such as pull requests, or even just diffs, will not be possible.
I submitted a feature request to the github folks years ago, but nothing has really happened (I was just suggested to delete and fork the repository again).
Not that it's hard: you could determine the ancestor and different lineages just using the hashes of the commits upon the first push to github. You could also do it completely offline, it wouldn't matter.
There's also quite a number of forks available on github which aren't really visible because of that. I know that for some of my own projects and smaller projects that I checked, a code search would actually reveal many non-linked repositories. And I also know why: I often don't fork on github (why would I if I know nothing about the project yet?), I just shallow clone locally. Forking on github doesn't serve any purpose until you actually change the code, which oftentimes has already been done locally.
git clone git@github.com:original/repo.git
# make some commits
git remote rename origin upstream
# fork the repo on github
git remote add origin git@github.com:your/repo.git
git fetch origin
git push -u origin masterThe main problem is not omitting the "fork the repo on github" part voluntarily. Sometimes you're just not aware that you're pushing a repository which already has some ancestor on github, while you originally cloned from the main author's website instead. This has happened to me countless times.
The graph network is totally useless in these cases.
One approach to this problem, as the article mentions, is to list by popularity - however what would this mean? If it's by the number of "stars", not many people curate their list to keep them up to date. It would have to be some kind of rolling popularity measure, perhaps number of unique users who've cloned a repository in the last month or something.
[1]: https://bitbucket.org/site/master/issue/5009/list-of-project...
git clone git@github.com:foo/bar.git
# fix locally
git commit
git push # creates pull requestGit evolved with a pull workflow because the problem it was made to solve was a the pull workflow of Linus and the kernel. This inherently means you must self host your changes while they're being reviewed and accepted.
As I wrote to an acquaintance earlier this week while venting about GitHub (and the condescending remarks you're liable to get from people who equate it with git and will assume that a tendency to stay off the former means you're unfamiliar with the latter):
"Coming from a background where wiki pages would be hosted on wikis and submitting [code] changes for review is as simple as a) creating a patch and b) attaching it for review, as I look at all the unnecessary (>3x) overhead that GitHub imposes and all the people who don't have a problem with it and feel that it's good and proper and normal, I feel like I'm in crazytown."
Further reading: Mozillians'comments on Gregory Szorc's post "Please Stop Using MQ"[1]. Pay particular attention to everything that Gijs has to say.
1. http://gregoryszorc.com/blog/2014/06/23/please-stop-using-mq...
It should feel like the inverse of the "check out pull requests locally" trick. https://help.github.com/articles/checking-out-pull-requests-...
It's harder to explain than to use actually. Ah! there's a bit of a wrinkle with gerrit in that it uses a local hook to insert an ID into commits, so rebasing or cherry picking knows which commit to reference. But that might be optional, it'd be like cherry-picking a pull request, I think github doesn't close the original in that case? Not sure on that though.
git clone git@github.com:foo/bar.git
# fix locally
git commit
git push # creates pull request
THAT ... would increase git user engagement by a factor of x>>0It's not as easy as DIY but it's just a support request.
In terms of not elevating the root to special status, I disagree. Recognizing one particular repository as canonical is a feature, not a bug. As much as git itself doesn't place any special significance on a particular repository, the culture of open source development does. Linus Torvalds can certainly say that his Linux repository doesn't hold special status, but that's just not true beyond a technical level. It's useful to be able to say "this is the main, supported version of this library"
First project pushed is canonical and can pass the baton?
Seems better to drop the idea of a root or canonical version totally. Linking forked projects together with hashes for the network graph seems like a good idea personally.
Isn't that what "branches" are for?
Oh, wait.
/s