We're doomed >_<
You would think it wouldn't be THAT hard to shard something like GitHub effectively.
I mean, all user accounts/repos starting with the letter 'a' go to the 'a' cluster and so on seems not exactly science-fiction levels of technology.
livejournal, facebook, twitter, linkedin, tumblr, pinterest all use (or formerly used) sharded mysql and most of these are at larger db size than github
i will also repeat my comment from another recent thread: i just cannot understand how 20+ former github db and infra people recently left to join a db sharding company. this makes no sense whatsoever in light of github's lack of successful sharding. wtf is going on in the tech world these days
I believe you have the chain of causality backwards here. In fact, I think it suggests that talent that went to planet scale is perhaps not the issue.
this is like if you were building a high-rise condo, would you hire the architects or management company from the building that collapsed in surfside florida? sure, they know what NOT to do next time, but that doesn't mean they do know what TO do
Platform migrations take a very long time and it's very complicated especially with decade old codesbases. I will say the current team at GitHub are nothing but outstanding people and engineers with a difficult task of managing a very large deployment.
Like, there’s another angle here: management, yeah?
Another way of reframing it is “maybe the folks hiring at planetscale know the inside baseball about GH infrastructure”. For example: https://www.linkedin.com/in/isamlambert
github should have sharded years ago, every other large mysql user did so much earlier in their growth trajectory
The hard part isn't solving the scaling problem, but the uncomfortable conversations with leadership about how chunks of roadmap will be drastically delayed. Inevitably, leadership makes a decision to kick the can down the road, further exacerbating the problem.
hundreds of engineers have worked deeply on sharded mysql at massive scale, many of us comment on HN!
but this is different. github is a 14 year old company, with annual revenue in the hundreds of millions USD
I literally cannot think of any other comparable size and age mysql user who has not successfully sharded long ago and avoided outages of this magnitude. and in return we get hand-wavy excuses of "complexity!!!" from their former vice president of engineering who was previously also their first DBA.
criticizing these excuses is not trivializing the complexity, it's more of a "we all did this at our respective companies, who also had a lot of complexity! why can't you? we are github users, we rely on you and are very unhappy, we want to know how this happened" and just getting excuses as an answer.
How about a little humility? You have no idea who else has similar problems out there. I'm sure Percona and other DB consulting companies would tell you otherwise, as would PlanetScale. (If this weren't true, they wouldn't be viable businesses.)
> github is a 14 year old company, with annual revenue in the hundreds of millions USD
You're making the classic mistake of thinking that having money means you can just wave your hand and hire whoever you want with whatever talent you need to solve your problems. The world just doesn't work that way.
As for the "hand-wavy" excuses, well, the people who know the issues don't need to disclose the dirty details in a public forum, nor are they often able to because of NDAs and other legal encumbrances. And it can be a career-limiting move to throw your colleagues under the bus.
You can choose to assume people are incompetent, or instead choose to assume that people are working as best they can under the constraints placed on them. I think it's better to assume the latter.
i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago, because that's what literally every other large comparable company has done.
there are front-page-of-HN threads about multi-hour github outages every single day this week. show me an example of another similar-size company having an equivalent meltdown please. only real equivalent is twitter during fail-whale days when they were only a few years old. they solved it early on, as should have been done.
i am not saying the entirety or even majority of github is incompetent, but i am saying there is clearly something extremely wrong and extremely unusual that led to this, and citing "complexity!" is just a pile of BS.
As I said, a lot of the problems organizations have are not publicly disclosed, and the people who know aren't authorized to talk about it. So I'm sorry, you're not going to get the examples you seek here. Companies like to keep their weaknesses close to their chest.
> i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago,
If not through money and hiring, how should they have accomplished it? You propose a goal but no actionable plan. That's worth diddly squat, both in engineering and in business.
if you think Percona and Pythian implement sharding solutions, you are deeply mistaken, this is not what they focus on at all. advice, sure. more on the perf and ops side. but not sharding implementation, and definitely not application side of things.
furthermore i am saying if other major companies were having this type of issue, everyone in the public would know about it! because the product/site/company would be down all the time. this isn't some thing you can hush hush. close to the chest? how do you keep daily outages close to the chest? completely absurd
there is literally no analog in the US tech world to a large 14 year old company having daily outages for a week, preceded by multiple outages per month consistently for the past year+.
i will stop replying now because it is clear we live in different realities or smth
I would also back otterley's position that there exists many other companies with larger non-sharded MySQL instances than github's. They may or may not have better reliability than github.
"They probably outsource the lock queuing logic to their Memcached layer." you are grasping at straws. the wrong straws. go attend some facebook conf talks, meet their engineers, read some arch papers, and stop speculating this nonsense. largest db tier there does not even use memcached, hasn't for 9 years. read the TAO paper!
horizontal sharding works great for many many companies with databases several orders of magnitude larger than github, why would it not work for github?
you are also saying there may be large companies with lower reliability than github? ok name them. again, if this was the case, everyone would know! their availability at this point is like what, 1 nine?
No, this performance cliff does still exist in the current mysql 8.0 branch. The Contention-Aware Transaction Scheduling added to 8.0 doesn't help the queuing problem with hot row at all. So facebook either doesn't suffer from the problem, e.g. all their writes are append-only so there is no hot row, or prevents the problem from happening, e.g. with API rate limiting built on top of something like memcached.
The improvement in innodb lock architecture mostly comes from splitting the locks into more granular levels. However, updates to the same hot row would require the same lock at the page-level. It is already as granular as it can be. So the tag of the bug report is correct, this is a problem of "lock queuing".
And frankly, given your attitude on this, I think you’re going to need to present some bona fides for anyone to take you seriously. Maybe you can even convince GitHub to hire you and solve their problems. I bet they’d be happy to listen to you berate them on how pathetic they are.
i have posted many correct technical details in various subthreads of this post. i don't care if you believe me, i do not live in your reality where constant downtime is totally fine and major companies are down all the time but no one knows about it
The Mythical Man Month has a few things to say about that.
(It's tempting to feel that the information is outdated, but in my experience it still seems true.)
This is an architectural problem, which even if they had the massive expensive brains behind something like mysql on their team they couldn't fix it.
(at least, I'm guessing, I think this kinda architecture doesn't scale even if they could kick the can down the road a few times..)
I think the complexity of GitHub’s data management requirements would surprise you. Better to withhold judgment until you possess all the facts.
this is classic Microsoft. spend a ton of money for something very valuable -- in this case virtually all developer marketshare -- and then casually pedal it into the ground while you lie about the KPI's to C levels (IIS marketshare on netcraft as a function of parked websites at GoDaddy to dominate over Apache) and keep it on life support with other revenue streams (XBox) for the next 16 quarters until it becomes a repulsive enough carbuncle to shareholders that it gets the axe (Microsoft phone.) then in a year, limp into the barn with another product nobody else but you could afford to buy (minecraft) and slowly turn it into a KPI farm for Microsoft account metrics to drive some other failing product (Azure) and keep the C level happy while you alienate virtually every player with mechanics or requirements they hate.
They don’t know Microsoft for anything other than ruining Minecraft. They didn’t know Microsoft made the Xbox or even Windows.
They made this statement after Microsoft forced them to migrate their account they’ve had for 5 years to a Microsoft account. That broke their computer for a few days and reset their games. For no useful reason.
galera has lower max writes/sec than a traditional async single master because it's a cluster. the other members of the cluster need to ack the writes, and all members are doing all the writes, so adding machines does not increase your max writes
I doubt they have enough to prove wrt SQL Server to make it worth going through that again.
If you meant that and I misunderstood, the other issue is that you don’t migrate to a different RDBMS. You rewrite half of the app and then spend a couple years fixing issues.
I would argue it’s more performant than vanilla MySQL and it supports multiple write masters using peer to peer transactional replication an enterprise license feature (https://docs.microsoft.com/en-us/sql/relational-databases/re...)
There are certainly some very niche use cases where you need it, but there’s a reason why tech stacks don’t use sqlserver. I think there are better ways to handle durable transactions and redundancy than using sql enterprise with sql‘s replication.
Nothing I have seen about Azure SQL fills me with confidence about their technical capabilities.
Performance at all tiers is woeful, the networking connectivity is madness, and disaster recovery doesn't.
"Ping" it using a trivial query such as "SELECT 1" a thousand times in a row and draw a histogram. Or just eyeball the numbers. I did this recently and had response times over 12 milliseconds regularly. For comparison, a small and cheap IaaS VM running SQL Server can get down to the 150 microsecond range and stay there.
For this 100x degradation in performance you get the privilege of paying several times the cost of an IaaS VM + SQL license.
Azure SQL proxies connections. I mean sure, the documentation says that they can do a "redirect" instead of a proxy, but not if you use any of their "private" network connection options. You would think that is some sort of simple IP header change implemented by the switching gear in hardware, but you'd be wrong -- it tunnels through what is essentially a VPN -- with all of the predictable issues. These tunnel VMs don't have accelerated networking turned on, for example. So you can have "business critical" tier on one end, and huge VMs with accelerated networking on the other end, but the traffic in between is being routed through some 1 vCPU virtual router appliance processing packets in software.
Anyone with database admin rights can alter firewall rules via SQL commands. These are via SQL Server accounts and hence they're just a username and password, no MFA or anything. If you can find a SQL injection vulnerability you can punch a hole through the "firewall". Brilliant.
The firewall supports IPv4 only, and uses different CIDR syntax to everything else in Azure. It doesn't support Service tags, or logging, or monitoring, or anything really. It exists only to tick a checkbox.
Unlike most other Azure services, SQL doesn't integrate with Azure Active Directory or RBAC properly. So for example you can have one AAD group as the SQL Admin. No other rights, no list of principal IDs... just one admin group. All other delegated permissions must be done through SQL commands, blocking the use of ARM templates, Policy, custom roles, etc...
If you delete an Azure SQL Server instance, it deletes all backups associated with it. Sure, they recommend that you put a "delete lock" on the server, but then most administrative operations become impossible because you can't then delete any child objects. And even if you do create a delete lock, that can just be deleted. There is no way to say "backups cannot be deleted for at least 'x' days", even though the underlying storage accounts support this feature.
DISCLAIMER: Some of the above may have changed since I last checked, always read the documentation and/or verify with support if your data matters to you.
That... is a massive footgun. Essentially a 1-click method to destroy your business.
I'll set up delete locks ASAP, thanks for the tip.
A DHT isn't a great fit for really rapidly changing data like metrics and such, but I figured that everything required to keep the basic git clone, commit, and web serving functionality is a pretty natural fit for a DHT.
I figured even code search had a relatively small amount of storage for indexing recent commits, with compaction of that data resulting in per-token skiplists stored into the DHT. That way, failure outside of the DHT serving path still allows code search for all commits older than the latest (incremental) compaction. Distributed refcounting or Bloom filters could be used for garbage collecting the immutable blobs in the DHT. The probability of a reference cycle in SHA-256 hashes of even quadrillions of immutable blobs is vanishingly small.
Some maintainer is acting in bad faith, someone else quickly locks them out of all the repositories that they have not yet thought to corrupt, then it looks really really bad if they find out that they still have permission to run the CI/CD scripts, maybe with malicious substitutions, even though you blocked them from being able to push commits up.
The other part of that is, it's not too hard to get right. See what GitHub was doing. People don't change ACLs that often, nowhere near as often as you read them. If they did, you could rate-limit them. And the rest of the problem is, you have one writer and many read copies in one place and you can enforce boundaries on how stale the data gets on the read copies.
Two hard problems, first is cache invalidation.
Of course, they have found out that they are big enough to shard, one would have hoped that they would have found that out in a gentler way, much sooner. But let's not pretend that their architecture makes no sense, it's a fine architecture, they just happened to outgrow it in a bad way.
If your CI/CD pipeline uses capability-based permissions, utilizing delegatable expiring cryptographic tokens (e.g. Google's Macaroons), then a stale read of the ACL (presenting a token incompatible with the latest ACL) would allow the CI/CD pipeline to start running. Completing all non-local I/O (i.e. observable side-effects) would require synchronous reads of a cryptographic hash of the latest repository ACL and comparison with an ACL hash in the token (with a slow path of re-verifying the token could still be issued in case of ACL changes). Presumably, this ACL hash would be a part of the repository's immutable root node, so in the common case of CI/CD pipelines storing outputs into a repository, this synchronous read is "free" in the sense that it's part of the transaction to change the reference to the repository's immutable root node.
Granted, you have to be careful and make sure all of the I/O that's not local to the CI/CD pipeline (i.e. all observable side-effects, e.g. releases, sending emails) goes through through an ACL check that involves a synchronized read of the repository root reference. The fast path is just checking the cryptographic hash of the ACL in the capability token against the hash in the repository's root node. The slow path (when the ACL has changed since the capability token was generated) involves re-checking the ACL and verifying that the user named in the token is still allowed to create the presented token.
The downside is a racy read of the ACL (in the form of presenting a cryptographic token based on a revoked permission) when starting the CI/CD pipeline results in wasteful resource usage. If this is a worry, an additional synchronous read can be added at pipeline start time to avoid this race. If the cost of this synchronous read is prohibitive, a cache of <user, revoked_permission, revocation_time> can be added to the CI/CD start logic.
The upsides are that you can cache ACL reads client-side (in the form of your cryptographic capability tokens), needing only synchronous ACL reads for non-local I/O (i.e. observable side-effects) and if an attacker starts a malicious CI/CD operation, locking out the attacker before the malicious I/O (releases, emails, etc.) completes results in the malicious non-local I/O (i.e. observable side-effects) getting discarded.
functional partitioning is a band-aid. you do it when your main cluster is exploding but you need to buy time. it ultimately is a very bad thing, because generally your whole site is depenedent on every single functional partition being up. it moves you from 1 single point of failure to N single points of failure!
I disagree, functional partitioning is not a band-aid, but an architectural changes that in the end can reap much more benefit than simple data sharding.
>> your whole site is dependent on every single functional partition being up. it moves you from 1 single point of failure to N single points of failure!
Not necessarily, it can also be that only some parts of your site are dead while others work perfectly fine.
then they fix it in post-mortem but pattern just repeats. i have seen it so many times! used to be much worse in the earlier days of the cloud when VMs would go poof more often
> In addition to vertical partitioning to move database tables, we also use horizontal partitioning (aka sharding). This allows us to split database tables across multiple clusters, enabling more sustainable growth.
Tech debt has to be paid in full, whether it's person-hours, or downtime, or both.