The way I see it is that there are two opposing things you can optimise for that depend on queue depth: utilisation or latency.
If you care about processing each message as quickly as possible then queues should be empty. This often requires a higher cost and lower utilisation as you inevitably have idle workers waiting for new messages.
If you care about utilisation, then you never want your queue to be empty while something is polling it. This might be some background task that runs on a GPU - every second that sits idle is wasted cash, and those tasks usually benefit heavily from batching inputs together.
In this case you want it to always have something to read from the queue and shut it down the moment this isn’t the case.