Building a high-scale chat server on Cloud Run
ahmet.im
ahmet.im
If you are a GCP or AWS or bare metal expert that can set this thing up in their sleep, that's great but the majority of people can really benefit from a PaaS like GCR.
Because Cloud Run uses vanilla Docker containers, once you have validated the idea you can move to GKE or VMs or a server under your desk or whatever. And if it never takes off, that's fine too because you didn't spend a ton of time investing in making it work.
Ahmet if you are reading this (hi) it would be REALLY cool to see something like using GKE for the base load and dynamically bursting to GCR to fill in the gaps. Not sure if this possible with GCLB today, but would be super cool.
(Disclaimer: I used to work on this team at Google.)
But if it “takes off” while you are caught with your pants down and you have no clue, you are stuffed with a friggin $60k bill and potentially have to remortgage your house.
And in the rare case it does, the infra is relatively portable (Docker container) and can be moved to a more-suitable VMs in less than the full month.
I think people underestimate the serving capabilities of a few VMs coupled with caching/cdn.
The cost of switching to something else if you actually need to is minimal. I improvised it a few weeks ago when I realized that doing background jobs wasn't something I should be doing on cloud run (it gets throttled). I added an endpoint that processed a data import in the background and that just wouldn't work in cloudrun. So, I wrote a script that runs the same container on a vm and then shuts it down. Took me about 10 minutes to write and it zipped through the import in a few minutes. For a proper setup, you'd want to use terraform or similar to provision the vms but that's not particularly challenging. Basically you point your provisioned infrastructure at your docker container and it will run it.
Very few businesses will ever see a quarter million concurrent users doing anything on their services. If you have that challenge, you can afford a few full time devops people to get you whatever infrastructure you need to handle this.
We've been running on cloud run since August. Standout features for me:
- it created a working CI/CD pipeline for me from my github repository with minimal effort and sane defaults. I got the service up and running within 20 minutes of me deciding to give it a try. I was so impressed with this that I decided on the spot to continue using it. That 20 minutes of devops time is technically the largest expense we've had so far on this setup as we mostly stay below the freemium threshold and it's otherwise a zero maintenance thing for us. Merge to master ... code deploys & runs. We've done this hundreds of times in the last months.
- comes with proper logging, monitoring, etc. tooling with absolutely no additional effort. I don't consider running blind an option so this stuff is usually not optional for me.
- we have no cloud run specifics in our github repository beyond the cloud build yml file. There is no vendor lock in. If you need a cheap way to run your docker container; this is it. Infrastructure is a blackbox that you get as a service. That's how infrastructure should be these days.
- once we got in to the test program; web sockets have been great as well. We just launched that feature a few days ago.
- low cost, low hassle, yet ticks all the boxes on the operational front. This is important to me as we don't have 24x7 people on staff and I hate having to deal with outages on a weekend or at night (as the CTO, I'm basically the person that gets to deal with this).
- I've been involved with devops on other projects. It can be a serious time sink. It certainly has sucked up months of my time. Many projects end up having several full time people dealing with that at typically premium consulting rates. Days of tinkering with terraform, chasing stackoverflow posts, reading piles and piles of documentation, etc. It just never ends. That stuff costs magnitudes more than you are likely to spend on hosting in your first year. Even the most expensive fancy AWS setup is a rounding error compared to that. Cloud run on the other hand is dead easy to setup and dirt cheap for small setups. It's great way to defer devops cost; potentially indefinitely.
> I realized that doing background jobs wasn't something I should be doing on cloud run (it gets throttled)
What do you mean by "it gets throttled". Do you mean that you exceed the CPU/memory limit in a container. Maybe you just have to plan the capacity for the background jobs: connections vs CPU vs memory. We use Cloud Run for background jobs with 0 issues.
In some cases we saw no evidence of the job even starting because the whole thing died right after the request returned. In other cases it would start but never finish. The actual import jobs weren't that big even. It was just super flaky.
The whole problem went away as soon as we gave it a dedicated vm. All I did was run the same container and then trigger the same request via a curl command sshed into the vm. Ugly but it got the job done.
2) it's 60k for a full month running at 100% capacity. Chances are you will notice this.
This blog post is showing the upper limits of GCR, not the typical use case.
"Oh no, we're too popular!". Who says this?
$60/k for your breadwinner seems like a deal.
30k was ingress and ETL, that amount didn’t include storage. Not sure what storage costs were off-hand.
> Also using GCR name for this is a bit weird when there’s already another GCP container product with the same exact name.
Very true! Naming is hard
I learned this recently and thought I'd share! From https://en.m.wikipedia.org/wiki/Amazon_Redshift
* BigQuery vs BigTable (though, there are at least some hints in these)
* Cloud Storage vs Filestore vs Datastore vs Firestore (vs Firebase? How are these related?)
I can of course make an effort to keep these straight, but then having to clarify in every design convo because my manager can't, or having a hard time googling (lol) things germane to the product I'm on, etc, is just a huge hassle.
Get CloudRun handling the C10k problem and you’ll really have something.
As my site matures, I might move some predictable parts to containerized VMs to save on costs. I have a crawler-type service that has extremely consistent traffic that probably shouldn't be on Cloud Run, but I did it anyway due to being able to iterate so quickly (see revisions here: https://imgur.com/fkFtIRM). Cloud Run also has a free tier which was enough when I was starting out and prototyping. It's also nice since my site gets very little traffic at night so at the moment the cost is not too bad. This is what my billing looks like: https://imgur.com/gIo3IGJ
If I was a big company or my site got super popular I might do things differently. But for a side project it has made me enjoy programming more than anything I've used before because I can focus 95% on code.
This is the (very consistent) traffic for my scraper that should probably not be on Cloud Run but it's too convenient for me to switch at the moment: https://imgur.com/RYLdp2t
This is the traffic for my web server, but since it's 30 days it's hard to see that within each day I get very large (relative) spikes in the middle of the day. I also had some huge spikes earlier due to some Reddit posts I made that are not shown in the graph. Being able to automatically scale to any load has been a life saver as I've tried to do marketing on social media. Otherwise I'd need to provision to handle as much as the maximum load I expect: https://imgur.com/RL534de
If you're a startup or need to handle random spikes from social media or other unpredictable sources, I think Cloud Run is worth a look. Another great use if for random jobs you might need to run every now and then. Web servers also seem like a good fit due to the fluctuations depending on the time of day.
I had a worker service running on Heroku. Very CPU intensive. The traffic pattern was extremely low throughout the day, but had completely unexpected surges.
On Heroku, my choices were: Paying $3k (basically paying for peak surge throughout the month) or having a lot of slow/failed responses during the surge.
Moved to Cloud Run very easily. Just a normal dockerized 12 factor app.
Now I pay ~$50/m and it automatically scales up when I need more workers.
If you want a managed app platform, it couldn't be more simple and cheap.
The scaling model of Cloud Run has been great. Now that they support websockets, I will be moving all my apps slowly to it.
My only gripe: I wish they had something for worker processes. I know that the rest of Google Cloud has solutions for it but being able to just spawn worker processes as part of the same deployment would be fantastic.
Have a look at GCP EventArc. It's a new way to trigger cloud run instances. You could create an EventArc trigger that wakes up a worker image, or some such. Basically,EventArc enables serverless event driven design across all GCP. https://cloudblog.withgoogle.com/topics/developers-practitio...
It may not fit your needs, but did you know that you can invoke other Cloud Run containers from a Cloud Run container?
Alternatively, you can enqueue Cloud Tasks from a Cloud Run container. Those Cloud Tasks can in turn invoke Cloud Run containers.
Not sure if they fit your usecase, but I use both of these approaches and they really expand the capabilities of Cloud Run for me.
Around €1500 will get you a second hand Dell R720 with 190+ GB RAM and 48 cores.
It would be cheaper to handle such scaling with a vm instead of functions. Any kernel with a bit of tweaking will easily handle that number of connections. One just needs RAM and sufficient CPU. 16GB RAM and 8 vCPUs on aws will be enough for 500k.
Sleeping would be very uncomfortable if my service using such approach started being popular and was running at increasing capacity for a week because my team needs time to move back to vms.
Ahmet makes that exact point:
> In the long term, as your load becomes more predictable, it makes more sense to move to VM-based compute (such as GCE or GKE) as several mid-size virtual machines can handle the same load, potentially 50x cheaper.
The shiny spot for something like Cloud Run -- or Knative -- is varying workloads. I described this as "workloads with unpredictable, latency-insensitive demand" in my book on the subject[0].
[0] https://livebook.manning.com/book/knative-in-action/chapter-...
Do people really consider putting a server app in a function?
Time elapses. What makes sense today may not tomorrow. It's fine to adapt designs when the economics point that way. Also fine to leave things as they are when the economics point that way.
It’s “interesting”
To prevent 404s hitting your function you have to define your routes in both your code and in api gateway. This starts to marry your code to your deployment.
Observability is now harder and you need yet another thing. We’re using X-ray with datadog.
To “safely” talk to Postgres you should use an rds proxy. It’s still unclear how much is different from pgbouncer or similar but in any case it’s pretty easy to knock over your db with runaway lambda invocations.
I could go on about the mistake we made of using SAM and how it ends up not supporting everything you need but I guess my main point is it’s easy to get sucked in by the idea of it being easy but you are trading that for service / operational complexity with a healthy dose of platform lock-in.
I need to spend some more time playing with Cloud Run because I wasn't aware of the concurrency bit you are mentioning.
I'd put a server app up behind a cloud run endpoint, sure, as long as it can fit in a few GBs and start in a few 100 ms. The scale-to-0 aspect is nice, and the scale up aspect is nice too.
Thanks, good to know it’s a common technique.
For GCP cloud functions, AWS lambdas, Azure cloud functions you must fight somehow against the environment (that's the reason frameworks like Zappa exist).
If you want/need to use compute instances, you can use KNative in your K8s cluster and enjoy a "similar" scaling experience.
Also happy client im easily handling hundred thousand connections for less than 80€/mo
Bare metal is cheaper, but you're paying 100% of the time for your peak traffic. If you only have a couple of peaks with 250k concurrent connections, with serverless you'd pay much less than €1500 but still be able to satisfy that demand.
Is it as simple as $60k/number of connections? If so the cross over point would be just 6250 connections as your base load. IMO, that $1500 server starts to look very appealing very quickly but I imagine it isn't as simple as that.
> Any kernel with a bit of tweaking will easily handle that number of connections.
I don't think tweaking a kernel and "easily" fit together. Unless you are kernel developer.
> I don't think tweaking a kernel and "easily" fit together. Unless you are kernel developer.
I think he means twiddling the kernel knobs in /proc to raise the limit of file descriptors. Last I checked, most distros default to 100k maxmimum file descriptors, with a different (10k) max per process.
Changing the max by 'tweaking' involved echoing a new number to a file in /proc.
epoll allows high concurrent client count, but how many requests per second and more importantly how many requests per watt!?
Also what is the response payload size?
I can do 1000 HTTP req/res (4K) / watt on my new atom 8-core server with my own HTTP app server passively cooled and tiny mini-itx size!
Web sockets are a handy thing to have access to but the app described is a pretty pathological case for the Cloud Run pricing model. I think people knocking the article on pricing alone are missing the point.
Given that, it seems like WebSockets/persistent connections is a weird use case.
Why do we now need thousands of instances to achieve the same?
Moreover, the 250K limit is an artificial limit imposed by Google for a generic cloud service. If you controlled _that_ software and tuned it for your specific application, you probably could squeeze an order of magnitude (or two or three) more connections out of it. Whether that's a worthwhile tradeoff is a function of whether the time required to rewrite the functionality of Cloud Run yourself (and run it, and maintain it) is lower than the cost of simply spinning up more boxes. If you're building a product and have the cash, it really doesn't matter whether you can run 250k connections or 2M connections if one could exist this week and suit your needs and the other could exist in 6-8 months when you build/scale/deploy it.
Engineering for highly variable capacity and engineering for extreme scales are two entirely different problem spaces. If a managed platform like Cloud Run makes your business more efficient and you can save on compute costs, the "efficiency" that you're talking about is completely moot. Being able to handle bursty loads when you need to and being able to consistently handle large volumes of traffic are two separate and unrelated problem spaces.
The second problem is that really the demo has just kicked your scaling problem over to redis. The demo isn't doing anything actually interesting, redis is doing all the work. The reality is that one of the best parts of Cloud Run over something like serverless functions is that you can have state in your server. You don't need redis at all if you are doing the same demo in Kubernetes without Cloud Run, since it's pretty easy to use channels in golang to do the same thing. However, in order for you to really make that work, you need some way to at the very least route the traffic consistently so that people watching the same resource get routed to the same server, or to be able to have cross node communication.
So this is a great start, but a little work to go before many people would be able to switch over to something like this due.
1) Egress is hand-waved, when in reality the free GBs would basically exhaust in a month on just pongs and reconnection requests, so doing nothing at all. That amount of people just sending single emojis as messages for just one day (to play along "short marketing event") is already matching monthly napkin estimate stated.
2) Isn't there's a good chance Memorystore wouldn't be able to handle this at fairly charitable load interpretations under that many CCUs? 1‰ of those 250k CCUs sending a message every second means Redis needs to publish 250k messages as well.
Again though, I'm aware it's only a hypothetical scenario demonstration, and it really does look cute in how it's simple to deploy. Fun stuff.
1) You’re right, 1 GiB free tier will exhaust fast. Beyond that standard GCP-wide egress rates apply. So if you were using VMs to serve this chatroom, you'd be paying the same https://cloud.google.com/network-tiers/pricing#premium-prici...
2) I think you’re missing out the part that in this setup only 1000 instances connect to Redis (using 1000 connections). We don’t initiate a per-user connection from the backend to Redis. Once a message is pushed to an app instance, the sample app pushes it to all its connected users.
> Any Cloud Run service, by default, can scale up to 1,000 instances. (However, by opening a support ticket, you can get this number elevated.) This means we can support 250,000 clients simultaneously without having to worry about infrastructure and scaling!
Aren't these statements at odds with one another?
I don't personally find the two statements at odds. If you have 250k connections 24/24 7/7 for a whole month, that should mean your business is generating more than enough money to cover the cost.
Or, at the very least, that the losses by not being able to scale are higher than the cost of the infrastructure.
If this is false and you'll be bankrupt by those 250k connections, you should obviously not let your instances to scale that much.
I think this is a far more valuable model than Function-as-a-service, where the runtime environment is strictly defined. Since “it’s just containers” you could write a web server in brainfvck if you felt like it, write a tiny dockerfile, then ‘git push origin/main’ and Cloud Run gives you a running endpoint. Okay, there is some devops work to set up that pipeline, but that is mostly boilerplate.
It’s Heroku but it scales down to zero (ie free) and up to 1000+ instances of your server in ms.
I remember that in 2012, the Rizon IRC network maxed out at 80,000 concurrent users on an AMD Bulldozer CPU.
API Gateway supports an unlimited (you need to ask for an increase from the 500 new connections per second rate) number of connections and is only around $0.25 per million minutes + $1 per million messages.
So just having 250,000 connections open would cost only around $270 per month (and scales up and down as you please).
1. 250 connections per container, and
2. 1,000 containers.
However the 250 concurrency limit does not refer to connections, it refers to requests. 250 concurrent requests can actually represent thousands of clients, depending on their think time.
https://cloud.google.com/run/docs/triggering/websockets says:
"Since Cloud Run supports concurrent connections (up to 250 per container)..." "WebSockets requests are treated as long-running HTTP requests in Cloud Run."