At this point it is a long time since I finished studying control-theory so I may be missing something obvious.
At this point it is a long time since I finished studying control-theory so I may be missing something obvious.
Control theory is mostly based on physical devices and analog variables, so it does not often fit with software systems.
However you are right nonetheless. Large networks contain a lot of caching of different kinds and other forms of data replication.
Caches warm up by transferring data and the available bandwidth to do so is never infinite. Especially when flipping traffic between whole datacenters.
Most load balancing systems are simply unaware of this.
It seems to me that control theory would be fantastic for setting and automatically adjusting some of these parameters. Distributing the same PID params and model to each client would ensure consistent/predictable behaviour as well.
P is for proportional - where you just multiply the error signal with a constant. What this does is to encourage the system to track the level of the input signal (but potentially with some lag)
I is for integral. This is where you integrate the difference signal over time. What this does is to reduce the lag between the input and output signal D is for derivative. This is where you feed back in the derivative of the error signal What this does is to damp down the swings in the system especially those that come from being too aggressive with the above two knobs.
Good controller design often comes down to picking the right weights for each of the three types of feedback functions you can input into the system. So in this example, it might be that you distribute your requests to servers based on how over or under loaded the servers are...
The integral and derivative part will tell you when spinning up more machines or when to tear down the one you already got!
> So in this example, it might be that you distribute your requests to servers based on how over or under loaded the servers are.
They had a working approach to distributing requests to servers, described at the beginning of the article. It's sessions (aka connections or channels in Google's literature) they're focusing on.
Responding also to siscia's grandchild comment:
> Or you can spin up more machines. The integral and derivative part will tell you when spinning up more machines or when to tear down the one you already got!
You're describing something like Google's autopilot, which I just linked to in another comment. A PID controller might make sense there, but it's only tangentially related to the problem described in this article.
You can control for the P90 latency and increase or decrease the number of connections to backend machines.
Similarly the backend can decide to drop connections if it see that the latency of the reply is too high, or if the CPU is too high or whatever other metrics make sense.
I don't see if a similar system would ever reach a stable-enough state.