I think the real issue is that without heavy content moderation it'll quickly get overrun with spam. A great example is to go find an abandoned subreddit and see what bad shape it's in.
I think the real issue is that without heavy content moderation it'll quickly get overrun with spam. A great example is to go find an abandoned subreddit and see what bad shape it's in.
With some of their service issues in the past few weeks (DC stuff that's been publicized), it could be less. To the very best of my knowledge no substantial resiliency work has been done recently.
I do know that a good number of folks on-call for core services were laid off while on-call, so, that bodes well. I feel bad for everyone left trying to keep things running.
One is that hardware fails and needs to be replaced. That requires people who know how to install the replacement hardware and deploy to it. That's assuming the new hardware is 100% compatible - that won't be the case for more than a handful of years.
Another is currently unknown security vulnerabilities, whether in their own code or in external packages they use. Those vulnerabilities are there and they will be discovered. Once they are, things start being taken down from the outside until the system collapses.
Yet another is bugs. Every system of this scale has a large number of bugs, many of them unknown. Some of those won't be discovered until the right conditions arise - the right combination of data, timing, etc. When they are finally triggered, some of those bugs will take down entire subsystems, some of which are critical to the product functioning.
There are many more examples like this. There is no such thing as indefinite resiliency for anything near this scale.
The modern version of this (kubernetes + AWS/GCP), if designed could likely continue to run for a long long time. Especially a product as simple as twitter.
Unlike your Solaris box, they are the target of constant advanced hacking attempts. I've been a part of the response when AWS was doing urgent work because of a security incident. The company I worked at was large enough to be paying AWS over a $1M a month when one such incident required dozens of our engineers working around the clock for three days to deal with AWS's response. We weren't even directly involved in the security issue. But without that engineering effort, our product would have shut down. There were other security incidents we were directly involved in and those would have taken us down without an even bigger response (whether or not we were running in AWS).
And then there are hardware failure rates. Hard drives alone fail at a rate of 1-2% per year[0]. Not a big deal on a single box. A very big deal when you have many thousands of hard drives - multiple drives fail every day. Unless you want to WAY over-allocate storage for redundancy. Even with that, there are surprising vulnerabilities to hardware failure at this scale.
----
* Live migrate won't upgrade the CPU family you're running on, so eventually someone/a something on your end will be forced to deal with migrating it, but that's O(years).
If development stops and the model is fixed bad actors can find the gaps given enough time