Really though, I think a lot of people use celery for offloading things like email sending and API calls which, IMHO, isn’t really worth the complexity (especially as SMTP is basically a queue anyway). Of course, YMMV depending on your use case.
However, I find it is often more worthwhile for:
1. Tasks which take a long time to run
2. Tasks which need to happen on a schedule, rather than in response to a user request.
There is an option 3 too, which is for inter-software communication. Eg events or RPCs, but I found Celery to be very much a square-peg-round-hole for this, which is why I developed Lightbus (http://lightbus.org). Lightbus also supports background tasks and scheduled tasks. /plug
Honestly, I think that mostly anything that doesn't depend directly in the current state of your infrastructure should be done asynchronously. I've had a lot of issues with systems that start up doing everything synchronously: you'll probably need to refactor it to be asynchronous in emergency mode during a crisis.
But anyway, how is your application supposed to respond after any of those failures? Is it just supposed to ignore the failure and thread on like if nothing happened, leaving your users on the dark? Is it supposed to reliably log every task so that it can retry anything that fails and in the worst case feed failures into some monitoring system/process? Or is it supposed to inform the user of any success or failure before the user can move on?
Queue software is only a good match for the first. For the second you will need to roll your own interface with the monitoring system anyway, so it's much easier to roll your own queues and get control of everything. The third one is best done synchronous, it doesn't matter the nature of the process or how long it takes. But funny thing is, I have never seen the first situation on the wild.
For example, sending a mail can take a while because mailservers have queues and whatnot. If you send a mail after signing up a user for them to verify their email, you don't have to wait until the mail is "really" sent before letting them know that their signup has been processed. You can tell a background worker to send the mail and return to the user much faster. For another example, in a previous job I worked for a big file sharing service. If user wanted their files deleted that caused all sorts of calls to AWS to actually delete the files, which could take a while. However, from the user perspective it was pretty fast because all we had to do was set the file in the database to "delete in progress" state and tell a background worker to delete the files. Then we could show the user that their files were being deleted within a couple dozen milliseconds instead of having to wait for all the AWS calls to complete.
The correct pattern should be a client submits a request to get something done to a thin layer, receives a ticket that allows it to claim the result and goes away to either check for the result via polling for the ticket or receives a call back.
Although this delay chain might be considered worth it you don't want to scale the frontend to multiple workers for some reason e.g. single-threaded evented runtime, or GIL runtime (python, ocaml), or if you want to avoid CPU-hard tasks being executed on your frontends.
In that case, it might be valuable to transform CPU waits into IO waits by moving the CPU work to a jobs queue, possibly running its workers on a different set of machines entirely.
It depends. If you needed a coordination layer or you needed to isolate certain types of traffic then it makes more sense. I assume your alternative here is "why not just have a new client tier doing work" which is a reasonable architecture too
> One aspect of this set up I’ve never been able to understand is how the application then gets the result from the worker?
Often they don't in these architectures. I found it a little strange that a job queue was being used to serve (what seems like) synchronous traffic. Usually I see job queues in the wild used for async/send-and-forget workloads
Personally I would rather chain multiple synchronous service calls if I needed a synchronous workflow. It's just simpler to me to stick web workers behind a load balancer and scale that. This is less elegant when things take a very long time though or are prone to retries. Either the client or server needs to be responsible for queuing/message persistence/retries - with services the client does it, with job queues the server does it
But as we decided to implement an MVP with a single HTTP request from the client, this whole separation doesn't make any sense, exactly as you noted.