GitHub’s success is largely due to the network effect and it’s entrenched status as the canonical code repository.
Besides libgit2, aka “the secret sauce”, is already open source. What are you waiting for?
GitHub’s success is largely due to the network effect and it’s entrenched status as the canonical code repository.
Besides libgit2, aka “the secret sauce”, is already open source. What are you waiting for?
They’ll be running it on a raspberry pi, not hosting the Rails or NodeJS repo.
Has it really ? I remember many threads here on complaining about Github's up time
https://news.ycombinator.com/item?id=5808496
I would describe GitHub's real "secret sauce" as the issue-tracking, wikis, project boards, and release management parts, that don't get represented in the repo itself.
Which is to say, if you wanted to commoditize GitHub (which is basically what "open-sourcing your secret sauce" means), you'd have to create some sort of library that allowed you to treat a git repo + all those other things as one structured data-object. You would be able to use said library to both operate on all those pieces of data locally; and to sync them between different Git hosting services that all share those features.
Or, better yet, figure out a way to put all those features into git itself, so that every git repo automatically transports those pieces of data alongside itself.
Git already has notes, and signed commits/signed tags, which are all those same kind of "objects that just happen to be there." So they don't need to copy Fossil's architecture; they can just copy the way that said object types interact as dependents of commits (while letting them get blown away when commits themselves do.)
If Git repos just "had" wikis, issues, etc. inside them, the lock-in wouldn't be there, so people would be switching between Git hosts all the time—and there wouldn't really be much value in a "git host" at all, beyond what just having a Git dir on your own server, plus a native-GUI Git client supporting the wiki/issues/etc. features, would get you.
I think though that 99%+ repos or there have zero issues, pull requests etc.
This is okay. It's just source control at the end of the day.
People clearly never cared about that, since Fossil ( https://www.fossil-scm.org/index.html/doc/trunk/www/index.wi... ) has these things and it never caught on
It's more than that. GitLab raised the bar here. Being able to run GitLab CE internally has de-risked the decision to test internally. For the next wave of customers in the space, familiarity with GitHub open source isn't enough.
If Github was made with micro services architecture, it could be split into open source "Client side + Test-backend" and closed sourced "Production-backend". Backend can be composed of some interface and multiple implementation such as test impl and prod impl. “the secret sauce” could be the "Production-backend" and MS have good in-house talents who operates Azure so no need for the help from OSS community to improve backend.
However for example +1 button took so long time to be implemented even though it seems like small change in client side code and some adjustments in database. That itch lead to the https://github.com/dear-github/dear-github open letter and signers listed here: https://docs.google.com/spreadsheets/d/1oGsg02jS-PnlIMJ3OlWI...
The last sentence of the open letter says: "Hopefully none of these are a surprise to you as we’ve told you them before. We’ve waited years now for progress on any of them. If GitHub were open source itself, we would be implementing these things ourselves as a community—we’re very good at that!"
I think that's OSS community want, including but not only I want.
I think you're drastically underestimating the amount of code Github is powered by and how freaking long any type of refactor/rewrite would take. We're talking about years.
Does anyone have link to any interview or article talking about granularity of their architecture?
It's talking about a component of Github backend called Resque (Distrubuted Job Queue)
As long as I see these thing the architecture is highly distributed and it's possibly composed of micro services.
I've been through a number of large rewrites/reworks that took monoliths much like Github (with many many many services behind it) and split them up into modular pieces and it's an insane amount of work that can take years. You simply need very good reasons (including business reasons) to do that.
Moreover, companies at these sizes just have a LOT of code all over the place. Tooling, infra, supporting services, etc... Not to mention it's just not useful to have external contributors for a business product like Github. Doing code reviews, addressing bugs that were introduced, spending time discussing things with contributors takes an incredible amount of time.
Basically if the reason you want Github open sourced (and reworked into some weird architecture you described) is so that people can contribute to fix things and add features....Github could/will just hire more devs to work on that.
Github.com = Github Core + Production Services and Infra
Github Enterprise = Github Core + Services need for Self Hosting
Maybe it's not worthy to proceed based on assumption but if there is something like "Github Core" which is shared codebase between prod and self-hosted, open sourcing the core can be an option?
Then possible path might be isolating least coupled (and small) components of client side code and open source?