Google Shows How To Scale Apps From Zero To One Million RPS, For $10
forbes.com
forbes.com
I know the goal of the test is to demonstrate how capacity can be scaled upward in very little time and with comparatively little effort. And for that this demonstration is impressive!
That said, the requests per second number is not especially impressive given that a modest single server can easily saturate gigabit Ethernet with 100-byte responses (and it's even easier to do so with larger responses) [2]. I am left wondering, does the $10 cost cover the cost of server instances and bandwidth? If so, that is a very good deal. The bandwidth charges for exceeding the capacity of a gigabit Ethernet connection (1M RPS with even the most trivial requests requires more than 1 gigabit) would be substantial with many hosting providers.
[1] https://gist.github.com/voellm/1370e09f7f394e3be724
[2] http://www.techempower.com/benchmarks/#section=data-r7&hw=i7...
To demonstrate scaling of the Compute Engine Load Balancing fanout we used 200 n1-standard-1’s Web Server running Apache v2.2.22 on Debian 7.1 Wheezy Images. Users are encouraged to use larger VM types for better single machine backend web serving, however here we demonstrated the scaling of the load balancer to backends and were not concerned with the backends themselves using every cycle to serve responses. Each backend web server received ~5K requests per second, which is an even distribution.
So, to match the peak rps of solid (but not top of the line) dedicated hardware appears to take upwards of ~120 instances of n1-standard-1 (assuming that it scales linearly, of course). Not a trivial number.
That said, I am impressed at how quickly this can scale up. If you have a site that normally runs fine on a couple of instances, but occasionally sees massive spikes in traffic, this could make sense. And from a purely engineering point of view, GCE and EC2 are quite interesting.
The goal was to measure the speed of scaling and load balancing vs egress. Bigger egress would not change the load balancing decisions.
Anthony F. Voellm Google Cloud Performance Engineering Manager @p3rfguy
I think I'm going to play with your script package over the long weekend in our dev cluster - Thanks! And nice work!
"PS... Cloud Performance is hiring :)"
Anthony F. Voellm Google Cloud Performance Engineering Manager @p3rfguy
Anthony F. Voellm Google Cloud Performance Engineering Manager @p3rfguy
Demand is elastic of course, and if you really want to scale in a cost-effective manner you also need to do auto-scaling. As far as I can tell (I have no direct experience with GAE), it's much easier on AWS. It would also be interesting to see if you can scale your pool of webservers from 1 -> 200 faster on AWS or GCE.
The article does quote @cloudpundit, who hits on the true point of the exercise: Relevance of GCE LB load test: With AWS ELB, if you expect a big spike load, must contact AWS support to have ELB pre-warmed to handle load. I would also guess that Amazon is working to improve ELB to behave similarly, especially now that Google's product has less restrictions than theirs.
1 min increment billing enables you to spin up massive Web clusters to handle spikes or Hadoop clusters so big that you can process the entire dataset it a few minutes, while paying less than you would for a smaller AWS cluster that's billed by the hour.
1m req/s is:
60 million+ reqs an hour
1.4 billion+ reqs a day
43 billion+ reqs a month
It would be a weird situation to be serving that much traffic but not being able to afford hosting.
It would be cool if someone could estimate the cost of what it would take to sustain 1m reqs/s with dedicated hardware or maybe an unmanaged VPS cluster.
This is an interesting marketing stragery. People who wish to launch a VM can choose CE and people who just want a sandbox quickly they can use App Engine? Though I am really skeptical about the future of App Engine if the CE became cheaper. I am sure if that happens, Google will do everything it can to migrate things over to CE. This is probably many years down the road...
I still think CE is really good for computations.
Compute Engine is for all the other stuff that you can't run on App Engine.
Incidentally, the datastore is now available as a stand-alone service, Google Cloud Datastore: https://developers.google.com/datastore/ This should benefit Compute Engine users.
The biggest drawback for App Engine is lack of async support. The only ways to scale are: multiple-threads (slow) or multiple instances (costly).
Most likely. But those aren't the only things developers care about when looking for PAAS.
There is that whole "Platform" aspect. And AWS destroys Google in this respect. It has far more offerings and more importantly it has a very large ecosystem of companies who will be in the same data centre who you can leverage e.g. Iron.io.
-Brian Head of Marketing, Google Cloud Platform
I know that there are api compatible systems out there, and one could roll your own (so to speak), it will get very interesting over time. I'm also curious what happens in the application space for docker.io based cloud offerings in the next year.
1 million requests per second is 20 times greater than the throughput in last year’s Eurovision Song Contest, which served 125 million users in Europe No, r/s != Gb/s.*
(Cue the down vote from the google up voters)
The GCE load balancer has none of these problems, which makes it a huge advantage over AWS and ELBs.
Disclaimer: I'm an engineer at Heroku. We manage dozens of ELBs for ourselves, and thousands of them for our customers.
Why?
Is 1,000,000 request per second a stupid, pointless number as no-one ever gets 1,000,000 requests per second or is this some sort of meaningful number?
As we saw when running the techempower benchmarks, simply going from the plaintext test to the single database query dropped the best performer from ~600,000 req/s to ~100,000 req/s. Throw in a bit more business logic, another query, and a slightly heavier response, and it is easy to imagine that 1 million req/s now sitting much nearer to 20,000 req/s.
My point being that, that 1 million req/s is a very optimistic number when used in such a comparison. Is it still an impressive max throughput? Yes. I just don't want anyone to think that they can now, say, host 50 netflixes on this setup.
Note: I realize you probably weren't meaning to directly compare those two numbers, but it somewhat read that way. I definitely do appreciate the context though - quite interesting to know that the netflix API was peaking at ~20,000 req/s in 2011.
The Google test is both a theoretical max throughput (that one wouldn't reach under basically any normal use case) and a test of the load balancer capabilities. The Netflix 20,000 req/s number is, instead, a real use case example.
My point was that one shouldn't directly compare those numbers and say, for example, that this GCE setup has 50x better throughput than Netflix.
I imagine that if Netflix were to stub all of their API calls with noops that returned 1 byte responses, they would be able to handle significantly more than 20,000 req/s. Basically, I don't think we actually disagree here.
So I downvoted you.
1 million requests per second is 20 times greater than the throughput in last year’s Eurovision Song Contest, which served 125 million users in Europe No, r/s != Gb/s
Not sure what the point you think you are making is here, but the Eurovision site was tested to 50,000 req/sec[1]. 50000 * 20 = 1MM
[1] http://googlecloudplatform.blogspot.com.au/2013/05/how-scalr...
Anthony F. Voellm Google Cloud Performance Engineering Manager @p3rfguy
GCE Load Balancing uses a single IP address, you can point your DNS there and forget about it.
What Anthony's post shows is that this IP address will be able to serve 1 (upwards?) of 1 million requests per second.
This matters because you have control over scaling your backend (design properly, add more instances), whereas you don't have full control over scaling your frontend.
Indeed, a problem we had during Eurovision was that mobile providers (it was a mobile app) would cache the IPs of our frontend Nginx servers, so scaling those wasn't as easy.
So this new GCLB essentially "solves" scaling your frontend. That's something I'd care about ;)
Hope this helps shed some light here!
What if I don't want to run your version(s) of CentOS/Debian, guys? Hmm? (I hope somebody from there reads this) :)
Still, its a very very interesting offering and something to keep an eye on. I bet nobody here is using it yet because its still early days and the thread devolves into such old classics like "you'll never be able to reach a human at Google" and "they changed up their pricing with Google App Engine that one time and my app was no longer free to run all of a sudden!". Sigh.
I personally think it would be great if you could spin up the same instances on either cloud and load-balance/fail-over as needed/cheapest. Docker makes it doubly exciting.
-Brian (@bgoldy)
Anthony F. Voellm Google Cloud Performance Engineering Manager @p3rfguy
This works out great and results in cost savings, as it can survive spikes, plus during the night the traffic is at most half or even less than the traffic you get during the day. I had a setup that was handling over 30,000 reqs/sec and during the night it kept about 6 h1.medium instances active, while during the day it could go upward to 20 instances, but was usually stable at around 14 instances.
This article mentions ELB, but I don't understand - does Google's Cloud Compute offer something similar? Can one vary the number of instances based on the incoming traffic or other metrics?
To get autoscaling, you can either roll your own — Google explains how here[0] —, or use a multicloud autoscaling solution (disclaimer: I work for such a provider).
[0]: https://cloud.google.com/resources/articles/auto-scaling-on-...
Now, judging from how busy the engineers are here, I'd say this is a bit harder than it seems!
Google has allowed us to run "Code" rather than worry about infrastructure. Because it is an "Engine" we don't have to configure much, and the autoscaling is fantastic.
Google Edge Cache means we get Better than CDN performance on static assets.
There are limitation to AppEngine, but because we can also leverage virtual servers we can create a hybrid environments that let us do things AppEngine won't like install C Libraries, or run Windows (we aren't doing that, but we could).
We have been very happy with Google App Engine and since we are running millions of pages a day through our Natural Language Engine it has worked out really well for us.
-Brandon Wirtz CTO Stremor.com
Most sites aren't remotely close to this artificial traffic pattern (1 packet request, 1 packet response).
It's kinda cool from an L4 load balancing perspective that it's only one fault tolerant IP address. In terms of L4 LB throughput though, a single box with IPVS will happily do 1M pps.