hundreds of engineers have worked deeply on sharded mysql at massive scale, many of us comment on HN!
but this is different. github is a 14 year old company, with annual revenue in the hundreds of millions USD
I literally cannot think of any other comparable size and age mysql user who has not successfully sharded long ago and avoided outages of this magnitude. and in return we get hand-wavy excuses of "complexity!!!" from their former vice president of engineering who was previously also their first DBA.
criticizing these excuses is not trivializing the complexity, it's more of a "we all did this at our respective companies, who also had a lot of complexity! why can't you? we are github users, we rely on you and are very unhappy, we want to know how this happened" and just getting excuses as an answer.
How about a little humility? You have no idea who else has similar problems out there. I'm sure Percona and other DB consulting companies would tell you otherwise, as would PlanetScale. (If this weren't true, they wouldn't be viable businesses.)
> github is a 14 year old company, with annual revenue in the hundreds of millions USD
You're making the classic mistake of thinking that having money means you can just wave your hand and hire whoever you want with whatever talent you need to solve your problems. The world just doesn't work that way.
As for the "hand-wavy" excuses, well, the people who know the issues don't need to disclose the dirty details in a public forum, nor are they often able to because of NDAs and other legal encumbrances. And it can be a career-limiting move to throw your colleagues under the bus.
You can choose to assume people are incompetent, or instead choose to assume that people are working as best they can under the constraints placed on them. I think it's better to assume the latter.
i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago, because that's what literally every other large comparable company has done.
there are front-page-of-HN threads about multi-hour github outages every single day this week. show me an example of another similar-size company having an equivalent meltdown please. only real equivalent is twitter during fail-whale days when they were only a few years old. they solved it early on, as should have been done.
i am not saying the entirety or even majority of github is incompetent, but i am saying there is clearly something extremely wrong and extremely unusual that led to this, and citing "complexity!" is just a pile of BS.
As I said, a lot of the problems organizations have are not publicly disclosed, and the people who know aren't authorized to talk about it. So I'm sorry, you're not going to get the examples you seek here. Companies like to keep their weaknesses close to their chest.
> i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago,
If not through money and hiring, how should they have accomplished it? You propose a goal but no actionable plan. That's worth diddly squat, both in engineering and in business.
if you think Percona and Pythian implement sharding solutions, you are deeply mistaken, this is not what they focus on at all. advice, sure. more on the perf and ops side. but not sharding implementation, and definitely not application side of things.
furthermore i am saying if other major companies were having this type of issue, everyone in the public would know about it! because the product/site/company would be down all the time. this isn't some thing you can hush hush. close to the chest? how do you keep daily outages close to the chest? completely absurd
there is literally no analog in the US tech world to a large 14 year old company having daily outages for a week, preceded by multiple outages per month consistently for the past year+.
i will stop replying now because it is clear we live in different realities or smth
I would also back otterley's position that there exists many other companies with larger non-sharded MySQL instances than github's. They may or may not have better reliability than github.
"They probably outsource the lock queuing logic to their Memcached layer." you are grasping at straws. the wrong straws. go attend some facebook conf talks, meet their engineers, read some arch papers, and stop speculating this nonsense. largest db tier there does not even use memcached, hasn't for 9 years. read the TAO paper!
horizontal sharding works great for many many companies with databases several orders of magnitude larger than github, why would it not work for github?
you are also saying there may be large companies with lower reliability than github? ok name them. again, if this was the case, everyone would know! their availability at this point is like what, 1 nine?
And frankly, given your attitude on this, I think you’re going to need to present some bona fides for anyone to take you seriously. Maybe you can even convince GitHub to hire you and solve their problems. I bet they’d be happy to listen to you berate them on how pathetic they are.
i have posted many correct technical details in various subthreads of this post. i don't care if you believe me, i do not live in your reality where constant downtime is totally fine and major companies are down all the time but no one knows about it
No, this performance cliff does still exist in the current mysql 8.0 branch. The Contention-Aware Transaction Scheduling added to 8.0 doesn't help the queuing problem with hot row at all. So facebook either doesn't suffer from the problem, e.g. all their writes are append-only so there is no hot row, or prevents the problem from happening, e.g. with API rate limiting built on top of something like memcached.
The improvement in innodb lock architecture mostly comes from splitting the locks into more granular levels. However, updates to the same hot row would require the same lock at the page-level. It is already as granular as it can be. So the tag of the bug report is correct, this is a problem of "lock queuing".
The hard part isn't solving the scaling problem, but the uncomfortable conversations with leadership about how chunks of roadmap will be drastically delayed. Inevitably, leadership makes a decision to kick the can down the road, further exacerbating the problem.