sorry for mean feedback, but is weird after seeing "transparency" as top value!
187 karma · joined October 26, 2016
sorry for mean feedback, but is weird after seeing "transparency" as top value!
i have posted many correct technical details in various subthreads of this post. i don't care if you believe me, i do not live in your reality where constant downtime is totally fine and major companies are down all the time but no one knows about it
"They probably outsource the lock queuing logic to their Memcached layer." you are grasping at straws. the wrong straws. go attend some facebook conf talks, meet their engineers, read some arch papers, and stop speculating this nonsense. largest db tier there does not even use memcached, hasn't for 9 years. read the TAO paper!
horizontal sharding works great for many many companies with databases several orders of magnitude larger than github, why would it not work for github?
you are also saying there may be large companies with lower reliability than github? ok name them. again, if this was the case, everyone would know! their availability at this point is like what, 1 nine?
if you think Percona and Pythian implement sharding solutions, you are deeply mistaken, this is not what they focus on at all. advice, sure. more on the perf and ops side. but not sharding implementation, and definitely not application side of things.
furthermore i am saying if other major companies were having this type of issue, everyone in the public would know about it! because the product/site/company would be down all the time. this isn't some thing you can hush hush. close to the chest? how do you keep daily outages close to the chest? completely absurd
there is literally no analog in the US tech world to a large 14 year old company having daily outages for a week, preceded by multiple outages per month consistently for the past year+.
i will stop replying now because it is clear we live in different realities or smth
i am NOT saying that github can hand wave this away with money and hiring. you misunderstand. what i am saying is company of their business size and scale should have already handled this long ago, because that's what literally every other large comparable company has done.
there are front-page-of-HN threads about multi-hour github outages every single day this week. show me an example of another similar-size company having an equivalent meltdown please. only real equivalent is twitter during fail-whale days when they were only a few years old. they solved it early on, as should have been done.
i am not saying the entirety or even majority of github is incompetent, but i am saying there is clearly something extremely wrong and extremely unusual that led to this, and citing "complexity!" is just a pile of BS.
but this is different. github is a 14 year old company, with annual revenue in the hundreds of millions USD
I literally cannot think of any other comparable size and age mysql user who has not successfully sharded long ago and avoided outages of this magnitude. and in return we get hand-wavy excuses of "complexity!!!" from their former vice president of engineering who was previously also their first DBA.
criticizing these excuses is not trivializing the complexity, it's more of a "we all did this at our respective companies, who also had a lot of complexity! why can't you? we are github users, we rely on you and are very unhappy, we want to know how this happened" and just getting excuses as an answer.
hundreds of engineers have worked deeply on sharded mysql at massive scale, many of us comment on HN!
then they fix it in post-mortem but pattern just repeats. i have seen it so many times! used to be much worse in the earlier days of the cloud when VMs would go poof more often
also their exact wording, "successfully", so if it fails your db is broken AND you do not get a shirt?
the person who for many years was directly responsible for github's databases, and later their overall infrastructure, recently made a horrendously inaccurate statement about the industry's most popular managed database offering
this shows either a severe lack of technical understanding, or a willingness to make deliberate technical misstatements
now github is down literally every single day due to database problems. it is very relevant! accountability is important!
they are encouraging people to record videos of making schema changes and reverting them, in order to win a t-shirt
does this seem unnecessarily risky or in really poor taste to anyone else?
this is not correct, for example pt-online-schema-change has long had a --reverse-triggers option which reverses the direction of the triggers to keep the old table up to date
github's outages have made my life difficult. planetscale's pricing scheme has also burned me. i am not just attacking this guy at random.
based on upvote points i can conclude many people do think i am contributing positively to this discussion. please refrain from tone-policing me. i don't see any technical comments from you in this discussion whatsoever so how exactly are your comments positively contributing?
I would have gladly had that discussion in the original context! but he did not reply to me there, nor to any of the other 6 people who replied to that horrendously incorrect comment of his
i would have let it go, but yet here he calls me out saying my own comment is grossly oversimplified, but he offered little explanation of his perspective. if i am understanding his brief argument it is:
1) github is old.
2) sharding is hard.
ok so here is my response: it was a lot less old when he started working there in 2013. many other companies have sharded several years after launch. yes it is hard. i have directly worked on teams sharding larger DB's than github's, sharding legacy apps. and so forth. it can be done.
it is really hard to justify github to be down this many times from the same root cause. year after year
i agree though my comment could have been nicer. apologies.
as to adding nothing to the thread: disagree. he literally dumped on AWS with false claims. how do you know i don't work at AWS? why is it ok for him to say stuff like that?
github should have sharded years ago, every other large mysql user did so much earlier in their growth trajectory
this is like if you were building a high-rise condo, would you hire the architects or management company from the building that collapsed in surfside florida? sure, they know what NOT to do next time, but that doesn't mean they do know what TO do
galera has lower max writes/sec than a traditional async single master because it's a cluster. the other members of the cluster need to ack the writes, and all members are doing all the writes, so adding machines does not increase your max writes
functional partitioning is a band-aid. you do it when your main cluster is exploding but you need to buy time. it ultimately is a very bad thing, because generally your whole site is depenedent on every single functional partition being up. it moves you from 1 single point of failure to N single points of failure!
livejournal, facebook, twitter, linkedin, tumblr, pinterest all use (or formerly used) sharded mysql and most of these are at larger db size than github
i will also repeat my comment from another recent thread: i just cannot understand how 20+ former github db and infra people recently left to join a db sharding company. this makes no sense whatsoever in light of github's lack of successful sharding. wtf is going on in the tech world these days
based on projections china's population peaked ~last year. it is a shrinking population from here, and while this will be a huge problem in most of the world (see https://news.ycombinator.com/item?id=30735230 discussed recently) but it is happening MUCH sooner in china, and at an unprecedented scale.
it is an existential threat and i am sure their government sees this, and it is scary to think what an "ends justify the means" way of thinking leads to with this problem
otherwise this way of thinking gets terrifying fast and rapidly descends into conspiracy theory land
example: "need to find a socially-acceptable solution to a demographic time bomb caused by decades of one child policy, while still maintaining ethnic homogeneity ? perform gain-of-function research to develop a vector that disproportionately harms the elderly"
to be absolutely clear, I don't believe this was actually the case in 2019 at all - but as an no-limits "end justifies the means" thought exercise - it is easy to arrive at inhuman dystopian nightmares
if they were all sharding experts why wasn't github sharded properly. other large mysql shops have solved this, all the way back to the days of yahoo and flickr and livejournal
if they had made github db/infra super-stable before this, it would be a vote of confidence in their new company, but instead imho it is the opposite
Not the same situation at all! With the US, "the" is officially part of the name of the country, used in the very first sentence of the US Constitution.
And even before that, in the Articles of Confederation, stated explicitly in its article 1: The Stile of this Confederacy shall be "The United States of America."
Why assume this is Western virtue signaling? The article explains the implications pretty well.
On the contrary, this is completely unclear because the mysql manual doesn't even give a clear explanation at all!! it only says "The number of rows read from InnoDB tables." https://dev.mysql.com/doc/refman/8.0/en/server-status-variab...
Let's say my db has a primary and 3 replicas. I insert one row. Does planetscale count this as 1 row written? 4 rows written?
Let's say my rows are small and there are 200 per DB page. I do a single-row select by primary key. This reads the page (you physically can't read just one "row" at the i/o level). Is this 1 row read or 200?
My read query triggers an internal temp table. Is the number of rows read based on only on the real tables, or also add in "reads" from the temp table? What if the temp table spills over to disk for a filesort?
I alter a table with 100 million rows. Is this 100 million rows read + 100 million rows written? or 0? something else? while this is happening, do my ongoing writes to the table count double for the shadow table?
Does planetscale even know the answers to these questions or are they just going by innodb status variables? do they actually have any innodb developers on staff or are they externalizing 100% of their actual "database" development to Oracle? (despite their blog post using wording like "we continue to build the best database on the planet"??)