Too Many Signals – Resque on Heroku
eng.joingrouper.com
eng.joingrouper.com
Solving this problem can be rather complex whenever Third-Party services are involved, but somehow this feels like you've only lowered the likelihood of multiple job executions, which isn't something I'd be comfortable with when it comes to things like credit card charges.
Discussion: https://github.com/resque/resque/issues/758
This is basically the best you can do. Imagine the context of a strongly consistent store you use for queues, and N workers. The store is strictly CP. One worker gets task "A", executes it, but does not provide the acknowledge, because it crashes before to ack. In order to guarantee at-least-once-execution property the system has to re-issue the task to another worker after a timeout, causing multiple job execution.
If you want to avoid this, you have as a side effect a failure mode where it is possible that jobs are lost forever, which is a lot worse, since the original problem can be solved by making all jobs idempotent.
Basically Redis-based queues should try to provide guarantees about durability of jobs, and try to provide just best-effort mechanisms to avoid duplication of jobs when possible.
Is it simply a case that this is the way that Heroku responds to being told to shutdown an instance? If so, why isn't the managing app that sends the shutdown call to the instance also handling the graceful "shutdown" of the processes on that instance?
* You want to instill the right culture in your customers code, that everything can fail, often, and they have to build software with that in mind. That's because stuff will fail always, and your customers will have to handle failures anyway, otherwise they will blame you.
* It's easier and cheaper to manage your fleet if all you have to care is that ninetysomething percent of your hosts are healthy.
* You can also detect broken machines more easily if you can remove it from the cluster as soon as you suspect it, knowing that no customers will be hurt. Where "broken" can mean anything, sometimes some instances will just run slowly, have bad IO, slow network, whatever, you don't care, you know you can just kill and respawn the containers as long as the total number of dynos meets the requirements and you don't exceed some predetermined rate of churn, which would affect the customer.
* You need to perform maintenance on machines where your run your customers containers/VMs. You can implement live migration, but it has a cost (implementation, management, storage etc), even more true a few years ago.
* You need to perform maintenance within the customers containers themselves; live migration won't help you with that. You don't want to bother your customers with maintenance windows.
* It's easy and cheap to "move" around containers across machines in order to balance load, spread an application across power domains.
I guess Heroku "dynos" are more suited for "worker" type jobs, then. In which case, sending the TERM signal to all processes isn't necessarily a really bad way of notifying the worker to shut down. Although, we are in 2014, and I don't see why they can't easily come up with a more robust solution. Even if it's in the form of a "shut-down" process, or giving the worker more than 10s to shutdown.
I sympathize. I've spent a heckuva lot of time getting clean shutdown working well (and someone just fixed a rare but persistent issue this morning!). There's a lot of edge cases. Steve and the Resque team are doing the right thing: you don't want a fix for one edge case to break another and this stuff is near impossible to test.
* Heroku sends the TERM signal.
* The process has 10 seconds to exit itself.
* After 10 seconds, the KILL signal is sent to terminate the process without notice.
Sidekiq does this: * Upon TERM, the job fetcher thread is halted immediately so no more work is started.
* Sidekiq waits 8 seconds for any busy Processors to finish their job.
* After 8 seconds, Sidekiq::Shutdown is raised on each busy Processor. The corresponding jobs are pushed back to Redis so they can be restarted later. This must be done within 2 seconds.
* Sidekiq exits or is KILLed.I noticed the "wait 5 seconds, and then a KILL signal if it has not quit" comment in the code above the new_kill_child method. Without jumping into the code, is the normal process sending a TERM, then forcing a KILL after 5 seconds? Just curious.
Yes, the situation you're describing is the RESQUE_TERM_TIMEOUT option which dictates how long the parent process waits to send a KILL signal after it send the TERM signal to the child. On Heroku you want that to be less than 10 seconds (and in practice more like 8 at max) otherwise heroku will terminate both processes with a KILL signal at the same time.
Obviously resque is closer to a "turnkey" solution and so forth, but what are the real fundamental differences?
The primary difference you'll notice is that RMQ has an explicit-ack mode. It will send a message to a client, the client processes it and sends an explicit ack (message consumed), at which point RMQ will send the next message. The client can also send a nack (push the job back onto the queue and redeliver it), and if the connection is dropped without the job being ack'd, then RMQ will requeue it and send it to another client.
If you're performing all your state mutations in a transaction or something similar that rolls back when a worker terminates, then you can avoid losing jobs and ending up in invalid state even during non-clean shutdowns.
As far as other notable changes go, you can have multi-queue routing (one message can be routed into multiple queues) and dead letter exchanges (so that TTL expired messages can be sent to a different queue rather than just being dropped). There's a lot more to it, as well; as a message queue, I do think that RMQ is flatly superior to Redis, but Redis has drop-dead simplicity going for it that is really nice if you don't need the extra features RMQ offers.
I'm happy to answer any questions you or anyone else might have.
RabbitMQ can be configured to not ack messages where an exception was raised, so if you have a durable store and the code responding to messages is idempotent/retriable ,you are good to go. Such a system can be easily configured with resque jobs using resque-retry, so it's mostly down to how you design your jobs/listeners/message handlers and not the underlying tech
I stopped reading right there, and thought to myself: Thank God I didn't choose Heroku as my service provider. Overpriced, and underpredictable.