Load balancing and its different types
wisdomgeek.com
wisdomgeek.com
This article only discusses web-based load balancing, which is absolutely important, but doesn't discuss supercomputer scheduling load-balancing. Its arguably a different subject... but the concept is the same.
When you have 4000 nodes on a supercomputer, how do you distribute the problem such that all the nodes have something to do? Supercomputer problems are sometimes predictable (ie: matrix-multiplications), and you can sometimes "perfectly load balance without communication".
But in the case of web-applications, there's probably no way to really predict the "cost of performance" before you start processing the service request (what if its a Facebook request to a really old photograph? Facebook may have to pull it out of long-term storage before it can service that request. There's no real way to know at the load-balancer whether a picture-request would be in the cache or not... at least, not before you process the request to begin with!)
-----------
In any case, I think adding "Predict the computational cost, calculate the costs you distributed to different nodes, and then distribute the new load to the node with lowest computational cost given so far" is a good method that works in some applications. (All blocks in a dense matrix multiplication have the same cost, so just keep passing out subblocks to all nodes as you're working on the problem)
Inherently, the idea that you're talking about boils down to having a way to characterize the nature of the request flows in such a manner that they can be evenly distributed. The ideal way to characterize them then, would be to know this information beforehand such that it does not require any computation at all to normalize the costs. As such, the best strategy would be to actually segregate traffic flows such that they're forwarded to "dumb" load balancers that use one of the strategies from TFA like weighted round robin.
Of course, there are many such optimizations available, but TFA seems to be targeting a beginner level introduction to a rather complex topic. As you describe, load balancing and scheduling algorithms have a pretty high overlap in terms of their theoretical foundations, and these concepts manifest themselves throughout any large scale system.
I worked at Facebook, but didn't touch anything related to this; so this is all conjecture.
The load balancer can't (or shouldn't) predict the cost to service a request; but the thing that generated the url for the image could; and that prediction could be passed in the url for the balancer to act on.
If you really need to balance by performance, it's probably simpler and accurate enough to provide frequent load feedback to the balancer. As long as you have a lot of requests, simple things work pretty well.
The vast majority of ISP caches won't keep your low TTL records in cache for years, but some do; this is a problem if you have to move your load balancers ever too though.
Depends on how stable your servers are vs your load balancers, and how many connections you need; and if you have enough IP addresses to give public IPs to your servers. Also, if you really absolutely need to control the load precisely, DNS isn't going to ever give you that.
Major cloud services like AWS support health/status checks through DNS these days: https://docs.aws.amazon.com/Route53/latest/DeveloperGuide/re...
It's also trivial to get around the client caching issue, just set a low TTL. Perhaps in the olden days providers had stricter limits on the minimum TTL you can set, but these days you can set it practically as low as you want.
EDIT: as a few commenters have fairly pointed out, TTL can easily be ignored by poorly-behaved ISPs and clients, so I'll admit calling it "trivial" to get around is not exactly accurate.
For the rare case of lift-and-shift-ing for a system upgrade I felt morally okay about eventually pulling the plug on them, but I'd hesitate to design a system that relied on well-behaved DNS clients if I had a reasonable alternative.
https://stackoverflow.com/questions/52032150/apache-force-dn...
Good to see ‘random’ redefined as pick random two then assign to one with least connections, which works strictly better than either random or least connections alone.
Misses a couple categories that may be relevant: least hops or best transit type network-mapped balancing to get to the ideal set of servers globally, as well as a technique that not-so-simply connects the user to the geography with the fastest response for them right then.
While the article notes value of fewer connections for a stream, you can take a bit longer in setup for a stream to get it right, as you and the viewer will pay the price longer if you get it wrong.
All this gets much more complicated when balancing very large objects, as you have to consider content availability and cache bin packing among the servers you balance to.
Load balancing isn’t just about server load or congestion, it’s also about network load and congestion. If a web page or video takes longer for a user to download, it ties up the server longer too.[1]
Load balancing algorithms can also consider network paths or round trip times between the user and a server to give users a faster web download or video stream. To do this, they may use information from network routing topology, such as how many “hops” or routers between the user and the server, or may even triangulate actual network performance by assessing measurements from multiple data centers and load balancing to the most responsive.
1. See “snoshy” comment on latency in these comments: https://news.ycombinator.com/item?id=25920284 — roughly, you aim to avoid queuing or connection creep, as you mentioned in the intro, and speed of opening, transmitting data over, then closing the connection, can make a huge difference.
You can write much of your site in raw HTML. For example, your right panel– you can write it in raw HTML, annotate it with some tags. Make it an iframe, just use CSS to get rid of the borders, or an XMLHTTPRequest. As long as you make good use of CSS, you can have a few tags and it'll work.
The downside is you lose any ability to balance based on details of the application protocol, it requires some specific network setup, and it's hard to find a DSR load balancer in managed hosting or cloud. I'm not sure if there's off the shelf software to manage DSR either (the basic pieces are there in most firewalls, but management isn't)