Load is not what you should balance: Introducing Prequal
usenix.org
usenix.org
This feels like one of those "my company realized savings greater than my entire career's expected compensation."
At the end of the day, what's left for the entrepreneur? You enjoy all the risk, but you don't get to have a paid 2 year sick leave. Even sympathy for your hard work can be hard to get. The potential of money is all you have.
Things are different in the US of course, where people can be fired the next minute without reason. That looks like just borderline abuse to me. But from a European perspective, the above comment does not deserve downvoting at all.
The Netherlands is the only country in the EU that taxes on unrealised gains.
Also worth noting that in the Madrid region they don't have the wealth tax.
So either 2 years of “stable income” or the chance of much higher compounding rate. As I see it, employment is a nice backup if being self-employed doesn’t work out.
Practically speaking, won't they fire someone by negotiating a "voluntary" severance agreement? It's not like an employee wants to stay on for long once it's been made known that they are unwanted.
Though obviously a severance agreement is better than at-will employment, it's also not a guarantee of long employment.
let's also not forget that the people involved didn't create this in a vacuum. it cost Google a LOT more than their compensation to make it possible for them to even start working on this project, let alone carrying it forward to completion.
people underestimate how hard and expensive it is to manage a company in a way that allows its employees to do a good job.
...let's also not forget that Google didn't manufacture this opportunity in a vacuum. It cost the rest of society a lot more than their revenue to make it possible for them to employ people who can spend their entire days working on software.
The actual "not a vacuum" context here is an environment that has been basically printing money for google for the last twenty years. It did not "cost" them anything. It's fine to acknowledge that the people who built google, however well paid they are, are creating vastly more value than they are personally receiving.
You say like it's something easy anyone can get done. Have you ever tried? If you try to build a sustainable and successful business, you'll see how hard it is.
Regarding infra, they were not capable of reaching the scale they needed and were already struggling by the time they got acquired.
On the money side, they couldn't attract advertisers as well as Google.
Given that YouTube was founded by 3 PayPal alums and was angel investor funded while being one of the top most visited sites at the time of its acquisition by Google, I am inclined to believe that their issues you mentioned above and others not mentioned would have been resolved in-house, via their partnership with Google Video or another company, or via acquisition by another company.
I don’t think it’s likely that YouTube would have failed as a company with the founders and investors it had, such as Sequoia, as it was simply too popular. YouTube had already started running ads for ~9 months before acquisition, though they were initially against pre-roll and other site-wide ads.
https://en.wikipedia.org/wiki/History_of_YouTube
> Participatory video ads were designed to link specific promotions to specific channels rather than advertising on the entire platform at once. When the ads were introduced, in August 2006, YouTube CEO Chad Hurley rejected the idea of expanding into areas of advertising seen as less user-friendly at the time, saying, "we think there are better ways for people to engage with brands than forcing them to watch a commercial before seeing content. You could ask anyone on the net if they enjoy that experience and they'd probably say no." However, YouTube began running in-video ads in August 2007, with preroll ads introduced in 2008.
You said “it cost Google a LOT more than their compensation to make it possible for them to even start working on this project, let alone carrying it forward to completion.” As if Google is groaning under the weight of sacrifices made specifically so these SREs can play in their engineering playground every day. I am saying this is backwards — that there has been no such sacrifice on Google’s part.
they're literally doing that. google has +$300B net positive assets. [1] they could distribute that to stockholders, go to the bahamas and call it a day. but they keep this wealth there so that SWEs can "play" and create stuff for others to use.
yes, they're not a humanitatian org. they do it so that they earn more, then invest more. but they are sacrificing short term gains and risking actually losing money.
That is not their motive or even a secondary consideration for shareholders.
If the fire department puts out a fire in your house you don't pay them the cost of the building. You don't give your life to a doctor, etc. That way of thinking is weird.
This contrasts with "willingness to accept", loosely the minimum compensation someone or some firm would accept to produce a good or service (or accept some negative thing).
Neither of these is sufficient to determine the price of something precisely, but, in aggregate, these concepts bound the market price for some good or service.
In my example of the fire department if nobody is really coming and you have no insurance or other way to save stuff you would indeed pay a lot. From what I read these were the dynamics in Roman times.
Often abbreviated somewhat sloppily into demand and supply.
The more requests you have, the higher the chance one of them hits a tail. So the overall latency a user sees is largely dependent on a) number of requests b) tail latency of each event.
This method improves the tail latency for ALL supported services, in a generic way. That's multiplicative impact across all services, from a user perspective.
Presumably, the number of requests is harder to reduce if they're all required for the business.
Indirectly, it is. As the quote I replied to suggests, in order to combat tail latency services often run with surplus capacity. This is just a fundamental tradeoff between the two variables mentioned.
So, by improving the LB algo, they (and anyone, really) can reduce the surplus needed to meet any specific SLO.
> Prequal has dramatically decreased tail latency, error rates, and resource
> use, enabling YouTube and other production systems at Google to run at much
> higher utilization.If you think about an object storage platform, much like with YouTube, traditional load balancing is a really bad fit. No two requests are even remotely the same in terms of resource requirements, duration etc.
With an Object Storage service like S3, no two GET or PUT requests an LB serves are really the same, or have the same impact. They use different amounts of bandwidth, pull different at different speeds, different latency, require different amounts of CPU for handling or checksumming etc. It didn't used to be too weird to find API servers that were bored stiff, while others were working hard, all while having approximately the same number of requests going to them.
Smartphones used to be a nightmare, especially with the number that would be on poor signal quality, and/or reaching internationally. Millions of live connections just sitting there slowly GETing or PUTing requests, using up precious connection resources on web servers, but not much CPU time.
However the "Replica selection" section seems to shed some detail, albeit somewhat indirectly. From what I can gather a probe consists of N metrics, which are gathered by the backend servers upon request from the load balancers.
In the paper they used two metrics, requests in flight (RIF) and measured latency for the most recent requests.
I assume the backend server maintains a RIF counter and a circular list of the last N requests, which it uses to compute the average latency of recent requests (so skipping old requests in the list presumably). They mention that responding to a probe should be fast and O(1).
At least that's my understanding after glossing through the paper.
a) are fast: they certainly incur the same network cost of a regular request. But more than that, all they do is read two counters, so they're super quick for backends to serve.
b) cheap: they don't do nearly as much work as a "real" request, so the cost of enabling this system is not prohibitive. They simply return two numbers. The probes don't compete with "real" requests for resources.
c) give the load balancer useful information: among all the metrics they could have returned from the backend, the ones they chose led to good prediction outcomes.
One could imagine playing with the metrics used, even using ML to select the best ones, and to adapt them dynamically based on workload and time period.
This appears a very logical solution, i.e., use estimation service quality instead of resource metrics, for scheduling. This is also more or less a known facts in the recent years, as systems are becoming so complex and so distributed intertwined that scheduling based on host load concerns a minor factors of serving requests. It's like one grew taller, and need to worry not stepping on huddles, but not bumping heads into door frame.
But we do need this kind of research to fomalize the practice, and get everyone on board.
Google's applied research absolutely winning here.
> PReQuaL does not balance CPU load, but instead selects servers according to estimated latency and active requests-in-flight
So, still load balancing
We present PReQuaL (Probing to Reduce Queuing and Latency), a load balancer for...
Fascinating that 2-3 probes per request is a sweet spot, intuitively it seems like a lot of overhead.
I really don’t understand how their claim is anything more than a least-conn with a better weighting algorithm.
We don’t generally use heterogenous server clusters anymore. Noisy neighbors and differences from one data center to the next are definitely things, but outside of microservices, you’ve got a lot of requests with different overhead to them. Route B might be five times as expensive as route A. So it’s not server predictors that I want, but route predictors. Those need a weight or cost estimator based on previous traffic.
Poor man version of this: we had ingress load balancers and then a local load balancer, like one does for Ruby or NodeJS or a handful of other languages. I found that we got much better tail latency running a more “square” arrangement. We initially had a little under 3 times as many boxes as cores per box, and I switched to the next biggest EC2 instance, which takes you to 3:4 ratio. That not only cancelled out a slight latency increase from moving to docker containers but also let me to reduce the cluster size by about 5% and still have a bit better p95 times.
I get two equally weighted attempts to balance the load fairly, instead of one and change.
Weighted round robin has some traction but even in engineering groups with lots of experience and talent the complexity of measuring CPU rate utilization is underappreciated.
Every production load balancer that I have come across in the last 10 years, load balances on the metric that is important to them.
NGINX appears to have Po2C available as an option on the "random" balancing algorithm. [1] It only considers load (requests in flight), unlike the paper which also considers recent latency (and does many other things).
However NGINX "Plus" (the paid version) appears to have a couple additional settings that use actual traffic as latency probes (not asynchronous synthetic probes as described in the paper). [also 1] It's not fully clear to me but I think if you use these options it's going to only use latency, not load. The paper described in this post actually uses both load (RIF) and latency. Eager to try these out.
Also, I noted that HAProxy author and kernel maintainer Willy Tarreau wrote an article in 2019 about adding Power-of-two-choices to HAProxy. [2] He builds a test bed with a few micro services and compares several balancing algorithms to Po2C, considering response time, throughput, and max load.
One conclusion is that "At the very least, Power of Two is always better than Random. So, we decided to change the default number of draws for the Random algorithm from one to two, matching Power of Two." However, in the response time data, the test appears to consider only average response time. I would love to see the p90 and p99 tail latency of these tests. The main focus of the paper in this post is minimizing p90/99, not average latency.
Tarreau also helpfully provides links to both a 2001 research report by Mitzenmacher, Richa & Sitaraman and also the original 1996 Po2C PhD thesis by Mitzenmacher.
[1] https://docs.nginx.com/nginx/admin-guide/load-balancer/http-...
[2] https://www.haproxy.com/blog/power-of-two-load-balancing
> weight wi is calculated as qi/ui, where qi and ui represent the recent query-per-second (QPS) rate and CPU utilization of replica i.
The pathological case describes a situation where antagonist loads are soaking up CPU on some on the machines, but WRR equally distributes the load because weight formula prioritizes the utilization of the individual tenant, not of the entire machine. Wouldn't including the the machine utilization (probably downweighted) into the WRR formula also solve the issue?
Funny to read this article today, just after myself and two others at work also just saved the company we work at far more than our combined expected lifetime gross earnings with a single (different) optimization.
I'm not done with the paper yet, but the basics are in fact written on page one.
Traditional load balancing has some known failure modes and engineers know to design to avoid them.