The single most important lesson from queuing theory for software systems is the non-linear relationship between utilisation and latency.
As system utilisation approaches 1.0 (100% capacity), the average waiting time does not scale linearly, it scales hyperbolically. A system running at 95% utilisation is vastly more fragile and slow than one running at 80%, even though the load difference is minor.
Think about a secretarial pool with five secretaries. The pool is at 95% utilization. This might mean that one of the secretaries is at 100% utilization and the other four are at 93.75%. It might mean that all of the secretaries are at 95%. It might mean that four of them are at 100% and one is at 75%. It might mean something else.
Going back to that third case, we can easily see that with the system at 95% utilization, at any given moment in time there is a 75% chance that the system is busy and unable to begin a new task, and a 25% chance that the system has capacity to begin a new task right away.
In the full-average case where all five secretaries are at 95% utilization, the odds that the system has spare capacity at any given moment are 22.62% (= 1 - 0.95 ^ 5), not too different from the case where four secretaries are kept busy 100% of the time.