Queueing Theory
en.wikipedia.org
en.wikipedia.org
- There aren't really any flashy results (the whole thing could be sold as restatements of "if the average processing rate is similar to the arrival rate, any variance in arrivals will lead to long queues".
- As a corollary, most queues encountered in practice can also be dealt with using very simple techniques or making the queue negligible. Neither of which require the study of queue theory.
- The quick recommendations are really boring (if you want the queue to get shorter, you need to process faster or add more servers).
But studying queue theory handy nonetheless because it turns out that queues are, in practical systems, about as common as list data structures or associative maps in programming. They are everywhere. Every time a stream meets a buffer, in fact. Being able to see a situation and reducing a lot of the noise to (M/M/n, lambda = 0.2, mu = 0.4) can free up a lot of thinking horsepower for more interesting problems. Then there is no need to try to reason about queue lengths vs serving times vs variance vs how those change with the addition of servers. An expert understands that a lot of results flow from a few simple variables, and doesn't have to remember the details because they are just symptoms of a few key observations.
So, in a sense, the reward for knowing a lot of queue theory is not having to think very much about queues.
This also applies to stuff like swapping / expiring cache - if you knew all of the future disk / memory accesses you could hit optimal, otherwise you need some heuristics.
It's been a couple of decades, but I think some the proofs involved playing the part of an omniscient adversary.
I don't remember the details now, but the proof wasn't too complicated, they gave it to us as an exercise.
I rather think that with "online algorithm", gnull means exactly this classic computer science topic:
Come to think of it, the obvious general principles of queueing theory have several consequences that aren't immediately intuitive to people:
- A system where concurrency is limited (practically all systems) often bottlenecks on its slowest component, meaning almost any upgrade will do nothing to improve its performance.
- The average time a task is stuck waiting is longer than the average waiting time, once there's significant variation (greater than Poisson) in arrivals.
- Based on only local measurements in an auxiliary, asynchronous component you can determine global throughput for the whole system.
The power is in the sheer number of semi-trivial observations that a queue theorist can start making after seeing only a small part of the system. And that is impressive - but mostly because you don't need to consider all those individual things as variables once the basic theory is understood. So the theorist can start ignoring all those variables really quickly and move on to dealing with the problem at hand.
1. Aperture automatically detects queue buildup based on metrics such as latency. 2. Adjusts the concurrency on a service. 3. Weighted Fair Scheduling of workloads (i.e. APIs) based on their labels.
[1] https://github.com/fluxninja/aperture/blob/main/blueprints/b...
Limit WIP!
I just spent a nightmarish time in Orlando. The theme parks there are insane, and I don't believe I'll ever go back. 2 of my friends and I literally stood in a two and a half hour queue for one ride and it cost us $150 each.
The ride was terrible, and the queue length was constantly being hidden. The companies have no incentive to increase throughput because building extremely long lines is cheaper than building efficient rides.
I’ve used them several times to show that a proposed system is mathematically impossible: “if the backend processes n requests/sec with a max response time of X, and the P95 of the backend is Y, the queue satisfaction is 0%”. People think you’re a wizard while you’re just plugging numbers into a century-old formula.
For a while for me everything was about queuing theory. Since then I keep forgetting and remembering things about this later.
Strikes me that I would like to have an easy to use and generic library for Rust with the various formulas etc. in it.
My experience with the formula's is a little different from yours. People thought what I showed them couldn't be true. "If the feed is X articles/second and system's P95 is Y, the backlog will continue to grow," would often be met with, "That can't be true ..."
Here's my website, if you want to see more of what I do: https://isaacg1.github.io/
Here's my question for you: queue theory is a nice tool, but most of the classic results are about systems (like M/M/c) that don't match the real world in important ways. Are there good rules of thumb for thinking about how different changes (e.g. burstier than Poisson, seasonality, constrained queue lengths, etc) map to these results?
Obviously, simulation is a powerful tool for more general systems, but being able to reason about effects quickly is super useful.
For thinking about bursty arrivals, a good rule of thumb is to look at the variance of the inter-arrival times. The key number is the variance of interarrival times divided by the mean interarrival time squared. The waiting time in a system with bursty arrivals will roughly be larger than the M/M/c by this multiplicative factor. Kingman's formula is the equivalent for the single-server setting: https://en.wikipedia.org/wiki/Kingman%27s_formula
For seasonality, if the arrival rates fluctuate over a long time period relative to the typical waiting time, it makes sense to just do separate calculations for the different conditions you experience. If the fluctuation is very fast, just use the average arrival rate.
For constrained queue lengths, there are a lot of theoretical results in this area, such as the M/M/c/c model: https://en.wikipedia.org/wiki/M/M/c_queue. The second "c" refers to the buffer size.
Thanks, that makes sense. More quantitatively, about where would get set the bar on "very fast"? Is it ~1x the mean interarrival time, or ~1 million x?
By the way, I really enjoyed your "Nudge" paper from last year. The result about FCFS was very surprising to me!
Here's a recent paper on the topic, if you're interested in the cutting edge research: http://www.cs.cmu.edu/afs/cs.cmu.edu/user/harchol/www/Papers...
Thanks, I'm glad to hear you liked the Nudge result!
One more: my understanding is that in Nudge I need to know processing time, but in FCFS I don't. How sensitive is your optimality result to errors in processing time estimates (I don't recall this being covered in the paper, but if it is feel free to tell me).
In the cloud services and databases settings, we seldom have accurate processing time estimates until we're quite far down processing a request (post-auth, post-parse, at least, but also for databases post-query-plan).
If the estimates were super noisy, you might be better off using Nudge very sparingly, only when you're more confident about the relative sizes.
First, many of these applications have a load-balancing step, where arriving jobs are dispatched to queues at each of the machines. The performance will depend on how this load-balancing is done.
Second, some applications will parallelize across many cores or machine, while others will run each job on a single core or machine. This obviously has major implications for performance.
Third, your application may have variance in interarrival times or in job completion lengths. This is also important to measure and incorporate into a model. Something like Kingman's formula can be useful: https://en.wikipedia.org/wiki/Kingman%27s_formula
Seven Insights into Queueing Theory [pdf] - https://news.ycombinator.com/item?id=20313834 - June 2019 (17 comments)
It’s Time for Some Queueing Theory - https://news.ycombinator.com/item?id=20290042 - June 2019 (62 comments)
Show HN: Queueing theory intro for software developers - https://news.ycombinator.com/item?id=18995551 - Jan 2019 (12 comments)
Queueing theory: The science of waiting in line - https://news.ycombinator.com/item?id=18983072 - Jan 2019 (134 comments)
Queueing Theory — Why the other lines always seem to move faster than yours - https://news.ycombinator.com/item?id=2038752 - Dec 2010 (15 comments)
But it turns out the simple parts of it go a long way! Eg at work we have a single deployment queue for a monorepo. At first approx this is an MD1 queue (deploy job takes roughly the same amount of time every time, though arrivals are actually way spikier than poisson), and realized that wait time was inversely proportional to total utilization.
While the infra team was saying “we do X deploys per day out of 4X and are only at 25% capacity”, I realized even hitting 2X would more than double the already bad wait time.
What happened is that a few initiatives were under way to increase capacity, but then out of nowhere the queue got log jammed (bc of high arrival rate variability) and we had to switch to Gitlab merge trains, which run CI concurrently on the optimistic result of merging. I wrote about it here: https://engineering.outschool.com/posts/doubling-deploys-git...
I’m planning on writing a blog post about the math of CI/CD deploy queues as G/D/1 queues.
For a programmer’s view of queuing theory and great performance testing foundations, I highly recommend “Analysing Computer system performance with Perl:PDQ” (don’t worry about the Perl, the book is very relevant) http://www.perfdynamics.com/iBook/ppa_new.html, which shows examples of queues inside computer systems and how to model them. The author has a nice little library to model different computer systems you come across.
I liked realizing that the dreaded “coordinated omission” problem in load testing (when you can only generate X rps and the server can handle more than that, your numbers are bunk) is actually when you think you are modeling an open system but you don’t have enough resources and end up seeing the closed system behaviour.
Why would you build a load generator like this? Normally because you run out of threads -- you have 800 parallel requests in flight and you can't open a new one until one has returned.
Correcting for CO takes a mathematical sleight of hand.
People who take the course by the author at CMU say very good things about the class and the professor.
But there are probably better books written since I was an undergrad.
If you want inspiration and some background, though, I can strongly recommend the book mentioned downthread: http://www.performancemodeling.org/
You're absolutely right that queueing theory would be helpful. A little part of me is wondering how on Earth you're doing your job without it!
1. I don’t understand how you “know” you’re in the steady state, beyond looking at the graph — I’ve always just made it go like 10k steps to force it, but this seems needlessly wasteful. I imagine you could track the derivative but if it oscillates?
2. I don’t know how to verify if the distributions are poisson/exponential in reality, except by plotting
More relevant note, I’ve always been surprised by how little exists online on the subject, and how little code it actually requires; I forget the terms but multiple event classes, multiple agent, infinite buffer sim took maybe 100 LoC C# from stdlib. Maybe 60 LoC in python — probably 25% was just tracking the stats I wanted like wait time.
Biggest performance trick is that poison inter-arrival rate is exponential, and so you can generate all arrival times up front; then instead of simulating every second, you can just skip forward in time to when the events actually occur… so your sim scales on number of events occurring, rather than length of time simulated. Pretty sure you can even do arbitrarily length simulations in constant memory but never tried
"How many operators I need to have at the same time when my telephone center receives 10 calls per minute with an average duration of 30 seconds so that the wait time won't be over 5 minutes."
The fact that such problems can be solved but I still need to wait 2 hours on the phone angers me very much.
Not understaffing, no no.
Exceptionally high call volumes all the time.
A: X, and then the operators will be busy Y% of their time.
Especially if Y is low, many callcenter operators will rethink what SLAs they want to give callers.
Many call centers optimise for keeping their operators busy aka paying as few operators as possible (and some of them then show their personnel the queue length or average waiting time and expect them to more rapidly kick out callers when that gets too long)
For most people, it’s the same with doctor visits. There is no inherent need to have to wait at a dentist, for example, but if dentists plan to have their day filled with paid work, there has to be. They may even schedule their first patient (possibly even the second) at a time before they plan to start working.
It’s different for people whose time is worth (much) more than that of the dentist. They (effectively) pay their dentist to wait for them.
[1] https://lists.rabbitmq.com/pipermail/rabbitmq-discuss/2010-M...
HN is for new information, not wikipedia pages that happen to be interesting but are already common knowledge. Although I guess it does technically fit the guidelines [1].
I've therefore always wondered whether there is software that uses theoretical results from queueing theory to improve control or to optimize parameters of complex systems. Any references?
(Arrivals still need to be Poisson, but that's a fairly nice requirement in that many real-life arrivals actually look Poisson.)