It's also a lot more expensive. Probably order of magnitude more expensive than the cost of a 1 day outage
It's also a lot more expensive. Probably order of magnitude more expensive than the cost of a 1 day outage
In theory the only additional cost should be the latency-based routing itself, which is $50/month. Other than that, you'll probably save money if you choose the right regions.
It is absolutely not as simple as spinning up a cluster on AWS at Roblox's scale.
Are the same instance sizes available in all regions?
Are there enough instances of the sizes you need?
Do you have reserved instances in the other region?
Are your increased quotas applied to all regions?
What region are your S3 assets in? Are you going to migrate those as well?
Is it acceptable for all user sessions to be terminated?
Have you load tested the other region?
How often are you going to test the region fail over? Yearly? Quarterly? With every code change?
What is the acceptable RTO and RPO with executives and board-members?
And all of that is without thinking about cache warming, database migration/mirror/replication, solr indexing (are you going to migrate the index or rebuild? Do you know how long it takes to rebuild your solr index?).
The startups you worked at probably had different needs the Roblox. I was the tech leach on a Rails app that was embedded in TurboTax and QuickBooks and was rendered on each TT screen transition and reading your comment in that context shows a lot of inexperience in large, production systems.
Seriously though, what is your RTO and RPO? We are talking systems that when they are down you are on the news. Systems where minutes of downtime are millions of dollars. I encourage you to setup some time with your CTO at Arist and talk through these questions.
2. I mentioned API Gateway and Lambda because OP asked if in general it is difficult to go multi-region (not specifically asking about Roblox), and most startups, and most companies in general, do not have the same technical requirements in terms of managing game state that Roblox has (and are web app based), and thus in general doing a series of load balancers + latency based routing or API Gateway + Lambda + latency based routing is good approach for most companies especially now with ala carte solutions like Ruby on Jets, serverless framework, etc. that will do all the work for you.
3. That said, I do think that we are on the verge of seeing a really strong viable serverless-style option for game servers in the next few years, and when that happens costs are going to go way way down because the execution context will live for the life of the game, and that's it. No need to over-provision. The only real technical limitation is the hard 15 minute execution time limit and mapping users to the correct running instance of the lambda. I have a side project where I'm working on resolving the first issue but I've resolved the second issue already by having the lambda initiate the connection to the clients directly to ensure they are all communicating with the same instance of the lambda. The first problem I plan to solve by pre-emptively spinning up a new lambda when time is about to run out and pre-negotiate all clients with the new lambda in advance before shifting control over to the new lambda. It's not done yet but I believe I can also solve the first issue with zero noticable lag or stuttering during the switch-over, so from a technical perspective, yes, I think serverless can be a panacea if you put in the effort to fully utilize it. If you're at the point where you're spinning up tens of thousands of servers that are doing something ephemeral that only needs to exist for 5-30 minutes, I think you're at the point where it's time to put in that effort.
4. I am in fact the CTO at Arist. You shouldn't assume people don't know what they're talking about just because they find the status quo of devops at [insert large gaming company here] a little bit antiquated. In particular, I think you're fighting a losing battle if you have to even think about what instance type is cheapest for X workload in Y year. That sounds like work that I'd rather engineer around with a solution that can handle any scale and do so as cheaply as possible even if I stop watching it for 6 months. You may say it's crazy, but an approach like this will completely eat your lunch if someone ever gets it working properly and suddenly can manage a Roblox-sized workload of game states without a devops team. Why settle for anything less?
5. Regarding the systems I work with -- we send ~50 million messages a day (at specific times per day, mostly all at once) and handle ~20 million user responses a day on behalf of more than 15% of the current roster of fortune 500 companies. In that case, going 100% lambda works great and scales well, for obvious reasons. This is nowhere near the scale Roblox deals with, but they also have a completely different problem (managing game state) than we do (ensuring arbitrarily large or small numbers of messages go out at exactly the right time based on tens of thousands of complex messaging schedules and course cadences)
Anyway, I'm quite aware devops at scale is hard -- I just find it puzzling when small orgs have it perfectly figured out (plenty of gaming startups with multi-region support) but a company on the NYSE is still treating us-east-1 or us-east-2 like the only region in existence. Bad look.
Also, still sounding like you don’t understand how large systems like Roblox/Twitter/Apple/Facebook/etc are designed, deployed, and maintained-which is fine; most people don’t–but saying they should just move to llamda shows inexperience in these systems. If it is "puzzling" to you, maybe there is something you are missing in your understanding of how these systems work.
You will note the problem was with the software provided and supported by Hashicorp.
18k servers is also not cheap, at all. They suggest at least some of their clusters are running on 64 cores, some on 128. I'm guessing they probably have a fair spread of cores.
Just to give a sense of cost, AWS's calculator estimates 18,0000 32 core instances would set you back $9m per month. That's just the EC2 cost, and assuming a lower core count is used by other components in the platform. 64 core would bump that to $18m. Per month. Doing nothing but sitting waiting ready. That's not considering network bandwidth costs, load balancers etc. etc.
When you're talking infrastructure on that scale, you have to contact cloud companies in advance, and work with them around capacity requirements, or you'll find you're barely started on provisioning and you won't find capacity available (you'll want to on that scale anyway because you'll get discounts but it's still going to be very expensive)
Not sure I agree. Yes, network costs are higher, but your overall costs may not be depending on how you architect. Independent services across AZs? Sure. You'll have multiples of your current costs. Deploying your clusters spanning AZs? Not that much - you'll pay for AZ traffic though.
Pretty much every city worldwide has at least one place providing power, cooling, racks and (optionally) network. You rent space for one or more servers, or you rent racks, or parts of a floor, or whole floors. You buy your own servers, and either install them yourself, or pay the datacentre staff to install them.
Multi-Region is a different story.
Data Transfer within the same AWS Region Data transferred "in" to and "out" from Amazon EC2, Amazon RDS, Amazon Redshift, Amazon DynamoDB Accelerator (DAX), and Amazon ElastiCache instances, Elastic Network Interfaces or VPC Peering connections across Availability Zones in the same AWS Region is charged at $0.01/GB in each direction.
https://aws.amazon.com/ec2/pricing/on-demand/#Data_Transfer_...
Wrong. Depends on the use case AWS can be very cheap.
> splitting amongst AZ's is of no additional cost.
Wrong.
" across Availability Zones in the same AWS Region is charged at $0.01/GB in each direction. Effectively, cross-AZ data transfer in AWS costs 2¢ per gigabyte and each gigabyte transferred counts as 2GB on the bill: once for sending and once for receiving."
Cross-AZ network traffic has charges associated with it. Inter-AZ network latency is higher than intra-AZ latency. And there are other limitations as well, such as EBS volumes being attachable only to an instance in the same AZ as the volume.
That said, AWS does recommend using multiple Availability Zones to improve overall availability and reduce Mean Time to Recovery (MTTR).
(I work for AWS. Opinions are my own and not necessarily those of my employer.)
In addition, unless you can cleanly survive an AZ going down, which can take a bunch more work in some cases, then being multi-AZ can actually reduce your availability by giving more things to fail.
AZs are a powerful tool but are not a no-brainer for applications at scale that are not designed for them, it is literally spreading your workload across multiple nearby data centers with a bit (or a lot) more tooling and services to help than if you were doing it in your own data centers.