I second this. [1][2][3] describe this sort of architecture. For you, the worker is a Mac. These articles are probably too long/detailed, because you don't yet need the availability/efficiency/elasticity they talk about. Your win is decoupling "handling the user request" from "doing the work".
The core of a worker is while (true) { download_job_request(); do_job(); send_result(); }. [4] Having the worker connect outwards to get the job, and submit the result, has benefits: (1) not having your worker "listen for incoming traffic" dramatically improves your security situation (2) you don't need to build a separate system to track which workers are accepting jobs right now.
I don't know your use case, or what physical and network security is truly needed. A Mac Mini or two in your house, ideally on a separate network [5], may be enough. Or you could rent from some cloud providers (Note: you do not want your Mac accepting connections from the Internet! You want a hardened remote-access method your cloud provider manages). Renting a cubic foot in a data center, and putting Mac Minis in it, is an in-between for cost/security, but will require the most expertise.
Workers running on cloud compute, or serverless, typically connect to the job queue using security permissions provisioned in that cloud. Just like your webapp talks to its database. If your workers are on-prem or in a different cloud, you'll need to provision access tokens/keys in your main cloud, and somehow make them available to your Mac software, which will use them to initialize the GCP/AWS/Azure SDK. Doing this "properly" is a big job! If you're going to YOLO and copy-paste them into a config file, I suggest (0) don't keep them in version control! (1) unique token per worker host (2) strictly limit tokens' permissions to what is needed (3) expire tokens after 120 days, meaning any worries about tokens issued in January, become moot by June.
You'd need a Mac-native solution to deploy and run software. If you have 3 "worker nodes", you can probably manage by hand. Past that, I've read Hashicorp's Nomad can manage deploying and running software on Macs like Kubernetes can on Linux.
You will eventually encounter situations like: (1) Workers lock up (2) Particular jobs take much longer than others (3) I just fixed an issue but there are 7 hours of jobs backed up in the queue... Until these cause unacceptable customer pain, I suggest you have one reliable and documented process to (1) Stop all workers (2) Empty the queue (3) Restart all workers. Practice this, to eliminate one source of stress during an outage. Also, make sure your UI behaves sensibly if the result for a particular job never appears.
If you get to the point of 24/7 monitoring, the things to page on are: (1) upper/lower thresholds for rate at which items enter the queue (2) same for items exiting the queue (3) upper threshold for queue size (4) upper threshold for per-job end-to-end time. All but the last should be trivial to monitor in a cloud queue product.
Definitely do some back-of-the-envelope math on the data size of the request and response. These can affect latency and cloud network egress cost. If that makes hybrid GCP/off-GCP impractical, I'd still recommend a worker-based solution that avoids exposing the Macs to user HTTP requests.
[1] https://docs.microsoft.com/en-us/azure/architecture/guide/ar...
[2] https://aws.amazon.com/blogs/compute/running-cost-effective-...
[3] https://shopify.engineering/high-availability-background-job...
[4] Note the Neural Engine is doing no work while waiting for downloads/uploads. You can get a lot more efficient by downloading the next job in the background, and sending results asynchronously. If macOS can handle a few different processes contending for the Neural Engine, you can likely get the same result by running multiple copies of a single-threaded worker per Mac.
[5] https://www.actiontec.com/blog/how-to-create-a-separate-wifi...