Rearchitecting GitHub Pages
githubengineering.com
githubengineering.com
There's a few differences. We don't use SQL in the routing chain, we use regex to pick out the site name and then serve from a directory of the same name (this is NOT as bad as it sounds, most filesystems can do this quite well now and take MUCH more than half a million sites to bottleneck).
DRBD is also a little hardcore for my tastes. Nothing wrong with it, I just don't know it well, and I don't like being dependant on things I don't know how to debug.
An alternative I wanted to show uses inotify, rsync and ssh combined into a simple replication daemon. It's obviously not as fast, but if you enable persistent SSH connections, it's not too bad. If it screws up, you can just run rsync. Rumor has it the Internet Archive uses an approach not too far away from this for Petabox. Check it out if you're looking for something a little more lightweight for real-time replication: https://code.google.com/p/lsyncd/
We're still working on open sourcing (!) our serving infrastructure, so eventually you will be able to see all of the code we use for this (sans the secrets and keys, of course). I've just been having trouble coming up with a good solution for doing this. For now, enjoy the source of our web app: https://github.com/neocities/neocities
Yep, this is basically our approach as well.
We've been using DRBD for quite a long time now on our Git fileservers (which also run in active/standby pairs - in fact, they look a lot like our Pages fileservers) so we have quite a lot of in-house experience with it and it's a technology we're pretty comfortable with. Given this, using it for the new Pages infrastructure was a pretty straight-forward decision.
This kind of thing is the way engineering should be. Kudos.
3ms for connection setup + auth + query seems reasonable. Are you using persistent DB connections? Any other mods? What sort of timeouts have you configured for DB connections?
I imagine the problems with DRBD mostly disappear if you're using it properly though, master/slave setups probably work really well.
I also am pretty conservative on engineering choices generally, and the "superfilesystems" (DRBD, Gluster) feel a little monolithic (read: not very unix) to me. It's not that they're bad, it's that they're solving a lot of hard problems, and there's a lot that can go wrong when you have to do that, and if something happens, you're the one that has to fix it.
I'm not religious about "do one thing and do it well", but SSH handles the transfers, rsync does efficient copying, and inotify fires events on file changes. Put them together and you've got a very "unix" solution. It's more or less an event-driven script that sits on the stable solutions to hard problems. If something goes wrong, you just run rsync.
I can't say enough how awesome OpenSSH is. I want to use it for pretty much everything. It's a work horse that really hauls.
https://kyledrake.neocities.org/misc/nginx.conf.txt
Critique away. As you can see, we've just barely avoided pulling out the lua scripting.
The next step for me would probably be to write something in node.js or Go. There's probably a lot of people cringing at that thought right now, but it's actually pretty good with this sort of work, and I'd really like to be able to do things like on-demand SSL registration and sending logs via a distributed message queue. Hacking nginx into doing this sort of thing has diminishing returns, we're kindof at the wall as-is.
I just want to applaud what you've been doing with neocites -- when the project started, I thought "Oh, nice." -- but not much more -- but I love the fact that you've kept at it, and your approach to openness is great (pending infrastructure code notwithstanding). I especially like your status-page:
(Which I found from your excellent update blog-post[1] -- but I think it could be even more discoverable. It's not linked from the donate/about pages?)
I hope your financial situation improves -- and still: I wonder how (almost) half-a-cent of revenue/month compares to most ad-funded startups sites? While you'll need... a "few" more to reach your goal. Actually "just" need 43x as many users with the same revenue/head to get there :-) )
ed: clarity (hopefully)
That's pretty amazing. If all you're serving is static assets, apparently you have to grow to pretty huge scale before one server will not be sufficient.
I'm curious if there was at least a caching layer, so every request didn't hit the SSD. They didn't mention it.
BTW, I'm often surprised how people are afraid from opening files from their scripts, as if they think this will always lead to disk access. Then, they start implementing a hand-written caching layer on top of that, which usually performs worse than what the OS already offers.
Also, BSD has some recent extensions to make sendfile work with ssl [https://people.freebsd.org/~rrs/asiabsd_2015_tls.pdf] (pdf)
We don't do any other caching on our own, although the other replies are correct in that the Linux kernel has its own filesystem cache which means not all requests end up hitting the SSD.
[0] https://help.github.com/articles/about-custom-domains-for-gi...
Fastly probably has something similar. You've really got to do this within 5 seconds or your user is going to get pissed off waiting for it every time they save/reload the page.
The big problem with passing caching to third party CDNs is that they need to be able to handle your SSL certs inline to request static files you don't need to change the URL of (because then you would be changing the published content which you don't control), and supports wildcards.
In effect you're doing what Cloudflare does. I'd use Cloudflare in a minute for this, but wildcards require their commercial plan and we can't afford it ($6k/mo). Also we would need fine-grained cache expire.
If Fastly does this and is priced in our range, maybe we should talk to them. :)
Disclaimer: I'm employed by Fastly :)
My rough estimates are that Neocities can handle 20-50 million sites with only two fileservers and a sharding strategy. So double it for each shard, and you've got a solution that actually scales pretty well.
1. Aren't static sites but are things like WordPress blogs running server-side PHP essentially unnecessarily.
2. Don't take down the server, but just hit the tiny bandwidth quota of some cheap shared hosting provider, so the host blocks it.
The fileserver tier consists of pairs of Dell R720s
running in active/standby configuration. Each pair is
largely similar to the single pair of machines that the
old Pages infrastructure ran on.
http://www.dell.com/us/business/p/poweredge-r720/pdTwo Xeon E5-2600s per machine.
For the most part, the load tends to come down to less than optimized database storage and queries on the backend, often duplicated for every page load for content that could easily be offloaded to the client browser, constructed client-side, and requested on demand.
It's a matter of striking a balance... In this case the content changes very infrequently, and having a publish step to static storage makes sense. It really depends on one's needs.
If you have a million pages you'd need to regenerate daily for a single site, and thousands every hour, then you might think differently.
Or is that automatically assumed when reading about a static hosting setup?
> We also have Fastly sitting in front of GitHub Pages caching all 200 responses. This helps minimise the availability impact of a total Pages router outage. Even in this worst case scenario, cached Pages sites are still online and unaffected.
The fact that a lot of non-technical employees in marketing and other fields are using it for corporate blogs is actually a nice bit of pressure on the organization to make Pages and web editing even simpler for those users. It becomes harder to lean on "oh it's a developer site so they'll figure it out".
Mostly, though, I think it's just a matter that we wanted it for ourselves. It's pretty awesome from an industry bystander's perspective to have something free, simple, and static, so we can all benefit from more stable docs, blogs, and so on. Maybe that'll change in the future and something Totally Different will change the industry, but for right now I think it's pretty rad, and totally worth the investment.
The CDN is key here, which you get if you use a CNAME (or ALIAS) instead of an A record for your custom domain on GH pages. I've found pairing pages with CloudFlare works great if you want to use a naked domain and you get HTTPS too. You can set up a page rule on CF to redirect all HTTP to HTTPS as well.
I'd pay extra for that, I (we all) have a bunch of personal sites, landing pages, marketing sites and tiny side projects that'd love to not have to deal with hosting – I think they'd make a killing, but also think must be in the works.
Github builds tools for developers. Atom, chat (abandoned), Pages, Gists, and github.com all fit within this. They tie into how teams operate. Serving JS is tangentially related — certainly something a web developer does — but not really core to their mission.
I think of this more as a convenience feature for their existing business that adds value. They use this instead of the "project page" design that the other code sites used, and it gives their users more control over the presentation. Which is awesome, and feeds into my conspiracy to get everybody to know and use HTML for presentation. :)
A few questions:
- Is everything in the same datacenter or in different datacenters? What happens if the datacenter is unavailable for some reason? Are data replicated somewhere?
- You moved from 2 machines to at least 10 (at least 2 load balancers, 2 front ends, 1 MySQL master, 1 MySQL slave and 2 pairs of fileservers). That's a lot more. Do you need more machines because you need more capacity (to serve the growing traffic) or just because the new architecture is more distributed and requires more machines by "definition"?
- I understand the standby fileservers are idle most of the time: reads go the active fileserver, only writes are replicated to the standby. Am I understanding correctly? If yes, it looks like a "wasted" capacity?
https://en.wikipedia.org/wiki/Cache_manifest_in_HTML5#Basics
[1]: https://github.com/jekyll/jekyll/blob/master/lib/jekyll/mime...
Edit: As you can see in that link, both .manifest and .appcache file extensions map to text/cache-manifest mime type.
And you could probably have a (small) hashmap that allows overrides for the most used pages.
Use hashing for the rest. Ideally consistent hashing, or rendezvous hashing, which i just read about on Wikipedia so must be good:
Say you have 8 servers at present, you'd allocate 8192 sequential keys to each server.
If you need more capacity, use 16 servers instead, and move half the files from each server to the corresponding new server.
Thus exactly half the data is moved each time the number of servers is doubled or halved.
[0] https://speakerdeck.com/jnewland/github-pages-on-riak-and-we...
A preliminary hit against a domain forwarder would be a good idea as well, but for those CNAME domains, dual-publishing might be a better idea... where the github name would be a pointer for said redirect.
While Cassandra itself might not be quite as comfortable as say mySQL, in my mind this would have been a much better fit... Replacing the file servers and the database servers with a Cassandra cluster... Any server would be able to talk to the cluster, and be able to resolve a response, with a reduced number of round trips and requests... though the gossip in Cassandra would probably balance/reduce some of that benefit.
Glad to hear it's being improved. I'm impressed that it was able to run on such simple infrastructure for so long.
The rest of GitHub runs on MySQL. As the rest of the post alludes, adding unnecessary complexity isn't what they are about (in this case.)
Basically, MySQL is a key-value store in an RDBMS's clothing.
Every time I've worked with mySQL I've seen some irksome behavior... just the same, setting up failover options is miles ahead of PostgreSQL. And when you already have in-house talent, it becomes even more obvious.
My only thought was that using a clustered database (such as Cassandra) as the store with the data itself might have been better. domain/url (minus querystring) would hash/distribute fairly well, and with even a relatively small cluster with 2 replica nodes for the shard would be pretty effective. Also, it would be easier to manage a replicated database, in my mind, than tracking sites to pairs of static servers. GoDaddy is/was moving to something similar with new development on one of their applications when I worked there, and able to serve a huge number of static requests (hundreds of thousands per second) off of a relative few servers with a sub 10ms response time, for content not backed by cdn.
In the end it just goes to show that serving static content on modern hardware can scale really well, with a number of options for technology. Which is why I'm somewhat surprised that something hasn't taken over the tide of poorly configured Wordpress blogs.
Just because most of the comments here have been (bizarrely) pro-MySQL, I have to link to this article:
So, if their team have lots of experience with MySQL but not so much with PostgreSQL, that could be a good reason to prefer one over another.
Have considered using cluster file systems such as GlusterFS or Ceph?
One story I heard from a PHP dev is that it would take 30 seconds to load a page while it looked for all the files needed to run it.
Here's the schema we use for the routing information:
CREATE TABLE `pages_routes` (
`id` int(11) NOT NULL AUTO_INCREMENT,
`user_id` int(11) NOT NULL,
`host` varchar(255) NOT NULL,
PRIMARY KEY (`id`),
UNIQUE KEY `index_pages_routes_on_user_id` (`user_id`)
);
Since we use MySQL for everything else, we decided it made the most sense to keep this routing data here rather than introducing a new database.Very smart. "Perfection is Achieved Not When There Is Nothing More to Add, But When There Is Nothing Left to Take Away"
I don't see anything about those components in this post. Did that architecture never make it to production?
And it doesn't work on custom domains (Github can't present a cert for your domain); you can make it "work" via Cloudflare but that only works if Cloudflare doesn't validate Github's cert, which introduces yet another unsecure link. [https://github.com/isaacs/github/issues/156]
We've talked a bit about our typical traffic in the past (https://github.com/blog/1992-eight-lessons-learned-hacking-o...), but suffice to say, we host some highly trafficked sites like the Bootstrap documentation.
If your blog isn't on that scale, and conforms to our terms of service (https://help.github.com/articles/github-terms-of-service/#g-...), we generally don't enforce hard usage quotas, but if you're concerned, feel free to reach out to support@github.com any time.
If these can be added to the free version, cool! :)
Full SSL support, integrated build automation (not just for Jekyll), redirects, built-in performance optimizations, rewrite rules, proxying, fine grained HTTP header control, integration with prerender services, etc, etc...
We can also do custom plans with true DDoS mitigation.
Why not use them to present customer projects (in private repos) to your customers? As long the pages are public this is pretty often a no-go.
Github pages is so cheap and easy it is the disruptive technology that is eating the lunch of stand alone hosting. Users don't think about servers, deployment or other things, they simply are pushing a branch and poof it is on the web.
Do you have a wishlist of features to add to GitHub pages? Maybe allowing minimal sandboxed server side computation with a max runtime of say 1ms, setting headers, redirects or other stuff? I am guessing every little addition would eat away at the alternatives.
Moving from WordPress, the fact that you couldn't constantly tweak a thousand little things was extremely liberating. That's the zen-like simplicity of GitHub Pages that to me, makes it an attractive option over heavyweight alternatives. Just push and your site is live. Fewer things to break and fewer things to worry about means more time to focus on what matters: your content.
/article/foo => /pub/article/foo.html
Or something very similar... That, or a custom 404 map, to take old urls, and send them to new ones (in the case of a blog, for example)... that said, gh-pages works well, and it's pretty awesome that it's offered to so many floss projects, and in my mind reduces the chances of a custom domain/website going away because someone doesn't really support something they put out there 7 years ago anymore.> Consistent hashing is a very simple solution to a common problem: how can you find a server in a distributed system to store or retrieve a value identified by a key, while at the same time being able to cope with server failures and network partitions? [2]
[1]: http://en.wikipedia.org/wiki/Consistent_hashing [2]: http://www.martinbroadhurst.com/Consistent-Hash-Ring.html
Pages also supports custom domains for both user sites and per-project sites, so we'd still need a way to resolve domains to users.
If you had a different business name...
I think a site that summarizes someone's github contributions from a recruiter / interviewer's perspective would be very helpful.
Poking around github to research a candidate is time consuming. Perhaps a more useful one page snapshot could be created. X profile views is free and you sell recruiters a per-company subscription fee.