Roblox has been down for days and it’s not because of Chipotle
theverge.com
theverge.com
Even though it’s a well documented issue with Postgres and you have an experienced team keeping an eye on it, a new write pattern could accelerate things into the danger zone quite quickly. At Notion we had a scary close call with this about a year ago that lead to us splitting a production DB over the weekend to avoid hard downtime.
Whatever the issue is, I’m wishing the engineers working on it all the best.
I think throwing together a static page better than "we're making the game more awesome" would be simple. It kinda makes me wonder if it's an internal auth/secret issue as has been speculated. That could theoretically make it harder to update the website, especially if it's deployed by CI/CD.
Quite surprising a seemingly battle-tested database can choke in such a manner.
in theory nowadays it wouldn't be too hard to change if you use logical replication to upgrade the database but it'd be a huge undertaking for a lot of companies.
This shows up sometimes also in numerical programming, e.g. when your meshing ends up producing more than ~4.2 billion grid points. But it is quite rare, precisely because UINT32_MAX is a fairly huge number.
If you’re wondering why it’s one: 1. it affects the size of a core data structure (that also gets serialized to disk and read back directly into said structure), and 2. basically nobody who has this problem (already a vanishingly small number of people) solves it by changing this constant, since you only have this problem at scale, and changing it would make your scale problems worse (and also make all your existing data unreadable by the new daemon.)
Also, FYI, Postgres uses compile-time constants for things that one might want to change much more frequently (though still in the “vanishingly unlikely” realm), e.g. WAL segment sizes. When you change internal constants like this, it’s expected that you’re changing them for every member of the cluster, or more likely building a new cluster from the ground up that is tuned in this strange way, and importing your data into it. “Doesn’t distribute well” never really enters into it.
With "doesn't distribute well" by the way I meant that it doesn't distribute well as program binaries, not across a cluster. It used to be extremely common to recompile e.g. your Linux kernel, nowadays almost nobody does that unless there are some very specific needs. Of course, building a specialized postgres cluster for exceptional scale would easily qualify.
I don't see why you should think that given that the discussion has been pretty much about the uint32 nature of TransactionId.
Reading their post history now they probably do know what they're talking about, though.
Most postgres databases wrap around with no issues.
The problem with increasing the size is a pair of xids are needed for every row, so you're doubling that if you go to 64 bits.
Speculation is a useful intellectual exercise, and the sign of a healthy, intelligent and curious mind!
Speculation is also fun!
I don't know how MySQL or MariaDB handles this; AFAIK it doesn't have this issue.
CockroachDB is an example of what a modern database should be like.
Because it's better than the alternatives? What would you suggest?
Even if we assume your premise that vacuum is at least as bad as the GIL is correct, this would easily explain it.
Add in the sheer difference in expectations between a programming language and a database... and well, I don't think there's any mystery here at all. There's easily multiple great explanations, even if your premise is correct.
There are companies that rely on Roblox for hosting their games so they can make money. This is the equivalent of cloud hosting going down.
My heart goes out to the Devs in this category, but hopefully this is further impetus not to stake too much on tech services.
I think to be working at the top of the Hierarchy of needs (in game development, for instance), people should demonstrably master the lower levels of the pyramid. We really can live in a way, where this is possible. We just have to collectively want it.
How much revenue would they have missed out on during this time? And how much will it affect their longer term growth?
If I were a Roblox shareholder, I would be pretty annoyed by the silence.
As a shareholder, you own the company including its problems. You don't get to complain like a customer, you did not buy any product or service from them, it's the other way around.
Huh? So shareholders who literally own a piece of the company don't have the right to complain about poor management, operational incompetence, company strategy, etc.?
That's not how it works. Shareholders actually have greater legal rights than customers even though both, in practice, are usually fairly limited when it comes to situations like this.
It's also a $50bn company, with thousands of shareholders and businesses that rely on their platform.
Would then posting every 2h "still working on it" really made it better?
[1]: https://en.wikipedia.org/wiki/Agar.io
[2]: "Around 2015 a multiplayer game, Agar.io, spawned many other games with a similar playstyle and .io domain". See https://en.wikipedia.org/wiki/.io
"STATUS UPDATE: Roblox is incrementally opening the website to groups of users and will continue to open up to more over the course of the day..."
- Perhaps internal systems they've developed, and the people who created them left. So it's not just fix the thing, but first understand what the thing is doing and then fix it. - Data recovery can take forever if you run into edge cases with your databases
Anyone found any articles about their architecture?
> 4 SREs managing Nomad, Consul, and Vault for 11,000+ nodes across 22 clusters, serving 420+ internal developers
> "We have people who are first-time system administrators deploying applications, building containers, maintaining Nomad. There is a guy on our team who worked in the IT help desk for eight years — just today he upgraded an entire cluster himself."
> A core system in our infrastructure became overwhelmed, prompted by a subtle bug in our backend service communications while under heavy load. This was not due to any peak in external traffic or any particular experience. Rather the failure was caused by the growth in the number of servers in our datacenters. The result was that most services at Roblox were unable to effectively communicate and deploy.
If they were to resume the game before restoring these issues, they would only exacerbate with state moving even further from where it was originally.
Or is this the highscore?
https://www.i-cio.com/management/insight/item/maersk-springi...
https://en.wikipedia.org/wiki/2011_PlayStation_Network_outag...
But I don't think they ever went down for two days?
But certainly Twitter were notorious for regular outages!
https://darknetdiaries.com/transcript/30/ "Shamoon" (the epsiode about this, which says it depends on the previous two episodes)
https://darknetdiaries.com/transcript/28/ "Unit 8200"
But if the organization is functional, in the medium term, this may also mean staffing understaffed teams, hiring SREs, etc. - which can mean less stress, no more 24/7 pager duty, better pay etc.
When something like this threatens to end the entire party (no one pays or get paid) you god damn want to figure out why and not make it happen again. That’s not burdensome, that’s business.
This is the burdensome part.
Perhaps (though IMHO not that likely) it may be some other kind of attack e.g. one intended to secretly steal customer data, which does not give signs of external intrusion if the company doesn't look for them much (and it might have motivation to not look very hard), but if it was ransomware, I'm quite sure they would not say what they said.
So while it's always a possibility, it's also kinda pointless to wonder about until there's supporting evidence.
I know they said they weren't hacked but they were hacked.
or
They are completely inept and have no disaster recovery plan in place, etc.
Best to them.
https://portworx.com/blog/architects-corner-roblox-runs-plat...