I'm a programmer, but I do mobile dev, and that amount of resources for a single machine sounds comically high to me. I thought modern websites usually had more 'normally' specced servers and just distributed the load across a ton of them?
I'm a programmer, but I do mobile dev, and that amount of resources for a single machine sounds comically high to me. I thought modern websites usually had more 'normally' specced servers and just distributed the load across a ton of them?
- 2 database servers (with half that RAM and 24 cores each)
- 2 redis servers
- 3 elasticsearch servers
- 11 web servers
- 4 load balancers
- 3 service workers
[1]https://nickcraver.com/blog/2016/03/29/stack-overflow-the-ha...Language/framework also matters. You're not going to pull this off with a backend based on a framework that gives no damns about performance. Rust, C# (bleeding edge), Java and C come to mind as good candidates.
I didn't keep the system in this state for very long, but for version 1.0 it was just the thing -- idea to production in a short period of time. Eventually it did move to a more distributed system, as log volume (and usefulness) increased, and we had more time to deal with the details. It was mostly nice to not have to reprocess data after releases -- I could do them in the middle of the day without anyone caring.
My biggest worry when writing this was that 40Gbps of network bandwidth wouldn't be enough, but it was fine in the early stages. 40Gbps is a lot of data.
I'm not sure I'd say it's a great sign that you need a single beefy machine to run something, but it's a tradeoff worth considering. I found the distributed system version easier to operate, but it did constrain what sort of features you could add. I think we got it right, but it's easy to code yourself into a corner when you have encoded assumptions deep into the system. Best to avoid that until you're sure your assumptions are right.
Avery wrote up some details of the system: https://apenwarr.ca/log/20190216 His idea, but I wrote most of the code ;)
There's a benefit to running on fewer servers. If you can make good use of many core machines and gobs of ram, it makes sense to go up to at least reasonably large machines. For Intel, dual socket Xeon is widely available and not obscenely priced; for AMD, I haven't seen a lot of dual socket Epyc, but 64 cores in a single socket is quite a lot. 768 GB seems big, but if you can put it into one machine instead of 12 machines with 64 GB, that helps reduce maintenance and communications overhead.
I ran systems with dual Intel Xeon 2690v4, a total of 28 cores/56 threads, and we put up to 768GB in some of them; that was several years ago, you can get a lot bigger now.
Databases love ram, and social sites make a lot of queries, so it makes sense a bit. I don't know what their usage numbers are, or what their site looks like; I'm just guessing based on general description. The traffic numbers didn't look too big, but types of request makes a big difference there; serving media is relatively easy, serving comments threads and highlighting your friends is trickier.
(Serving media with transcoding is a lot less easy though)
Your website's portal is down, I tried tweeting and email but I'm not getting a response. Just wanna make sure everything is ok