Since we don't have infinite space, we can expect eventually to lose some messages in this scenario.
Since we don't have infinite space, we can expect eventually to lose some messages in this scenario.
This is literally the slippery slope fallacy. You aren't accounting for minor fluctuations or finite timescales. If you were driving a car, you might say "if you aren't steering into oncoming traffic then you're steering towards the ditch" and conclude that no cars will ever safely reach their destination.
If the only valid state for a queue is empty, then why waste time implementing a queue?
If you feel like you have enough spare capacity and any given item isn’t taking too long to process then it doesn’t matter if you are ever empty.
Corner conditions are where I start to care:
* What is the worst case latency? How long until the work can be done? How long until the work can finish from the point at which it enters the queue?
* What is the worst case number of items in the queue? How large does the queue need to be?
* What is the maximum utilization over some time period with some granularity? How spiky is the workload? A queue is a rate-matching tool. If your input rate exactly equals your output rate at all times, you don't need a queue.
* What is the minimum utilization over some time period? Another view on how spiky the workload is.
Minimums and maximums I find are much more illustrative than averages or even mean-squared-errors. Minimums and maximums bound performance; an average doesn't tell you where your boundary conditions are.
In general you don't want your queue to fully fill up, but like the other poster said, it's some tradeoff between utilization and latency, and the two are diametrically opposed.
The unifying aspects of using queues for us is that it allows us to load-balance jobs across workers, and allows us to monitor throughput in a centralised fashion. We can set alerts on the time-sensitive queues and then use different thresholds for the background queues, but we're using the same metric data and same alerting system.
If there is a 1 hour queue for the checkout at a grocery store, you probably won't join the queue - you'll go to a different store...
[1] https://en.m.wikipedia.org/wiki/Little%27s_law
[2] broadly construed. For example a service that gets a mass of requests at the top of every hour displays seasonality.
Sure, that's technically correct but applies to basically everything. It is very likely your users/orders/whatever table in your database is also "trending to infinity" over time, except this doesn't mean databases should always be empty or that we should expect to eventually lose some users/orders/whatever.
Or, more succinctly, if your queue capacity is 10 million messages and your queue messages represent "houses purchased through our website", then in reality your queue capacity is infinite because nobody will ever purchase 10 million homes through your application per small unit of time.
But this is also where the disagreement comes in with flexible serverless workers. They generally do not cost more, so you'll end up spending $24 in an hour rather than $1 an hour all day. If the serverless costs more then you are likely not working through your backlog.
I agree it should be a business decision, so you catch up on weekends instead for example, but non-technical folks don't always grasp those numbers and that they are getting a lag with their data.
It depends what you mean by "approaching empty". If you mean "empty most of the time", then no, that's wasted resources. If you mean "its size is decreasing most of the time", then no, not possible. Even if you mean "it is empty some of the time", then maybe, but that is not a strong requirement either.
Of course more average input than average output is bad, but stating it as "should be empty" or "should be approaching empty" seem to suggest the wrong thing entirely. It should be smaller than the target latency of the process.
70-75% is just a rough number, the real number depends highly on several factors related to your specific use of the queue.
It may not make sense but it is the reality.
So I disagree entirely with "it's either shrinking or growing". It is both growing, shrinking, and stable, at the same time, over different horizons, and I have no idea what law you are trying to state.
The link you provided above is a really good primer based on serious math, I don't understand how you think it supports your vague claims.
Queues are either growing (adding items) or getting smaller (consuming items). There is not a state where they stay the same length, unless you're doing something really weird like monitoring to make sure there are always 10 items in the queue, or something, and adding if that number drops, but then you'd need a second queue-type structure to support that, and the conversation kind of goes off the rails.
I literally cannot understand how you think time works in your area.
You're also being really rude for no good reason. My polite suggestion is that you reflect on yourself here about why you are being rude on the internet to strangers. It could lead to growth and hopefully help you with whatever hurt you're experiencing.
I think the reason you get the responses you do is people are used to systems with high variation in flow rates that can only process a limited number of items at the same time, and for these, queues can soak up a lot of variability by growing long and then being worked off quickly again.
In such a case, you'd expect at least N messages in the queue at all times — and often a small multiple of N. Not because the consumer can't consume those messages in a timely manner, but because, given a reliable queue, messages can't be consumed until they've been 2PCed into the queue, and that means that there are always going to be some messages sitting around in the "acknowledged to the producer, but not yet available to the consumer" state. Which still takes up resources of the queueing system, and so should still be modelled as the messages being in the queue.