Lots of ways to skin this cat...
Lots of ways to skin this cat...
As for the SPOF: my logic was that if the DB is dead, there's no point having the ZeroMQ components come up anyhow. The jobs will keep on trying to come up, but until Ops brings the database back up, they can't do any damage.
It's amazing how small use case differences can mean very big differences in the effectiveness of various strategies. The at least once / at most once / no guarantee either way but soft real time, use cases, can mean radically different toolchains. This is the real lesson of distributed streaming.
I think it's much less complex than running an entire database node just for this, which btw will also require you constantly to poll, and will require you to bring in an (often heavyweight) client library into each node too, as opposed to standard-library sockets which if you're running multinode you're almost certainly already importing. If you're looking for "simple" distributed computing my sense is that that has yet to be invented.