Show HN: Workq – Job Server in Go
github.com
github.com
There is also a bit of oddness experienced when people try to use the various AWS libraries to hit non-AWS hosted endpoints since the happy path for these is to connect to AWS regions.
I worked on an internal SQS clone in the past (not related to current employer).
From a queue API perspective, I like iron.io's MQ API [0] and Google Cloud PubSub[1].
Disclaimer: I currently work for Nest, an Alphabet company
[0] http://dev.iron.io/mq/3/ [1] https://cloud.google.com/pubsub/reference/rest/
From going through the source, it looks like the payload is the cmd. Can those be multiple words or will it read each word as a separate argument.
handler := s.Router.Handler(cmd.Name)
reply, err := handler.Exec(cmd)
The ping example was just given so I could enqueue a job named “ping” as a client and respond with a “pong” text result from a worker. Just a silly example :). Real use cases would be background jobs such as sending emails, http downloads, image resizing…etc.
"ping" here is just a message, not the name of a system command/executable - in this case the worker receives the "ping" request, and just replies with a "pong" message - both sides are code you need to write.
You can encode whatever you want in the payload, json, simple plaintext - it's up to the client and the worker to agree on the meaning.
The biggest limitation of beanstalkd IMO is the fact that robustness features have to be handled by the client. That's why we're considering switching to disque or alternatives. It doesn't look like workq supports this, and it's likely something that should be designed in from the ground floor.
But beanstalkd's proven history is a huge mark in its favour that we'd be loathe to give up.
We're building an in house CI system (we have some weird requirements and can't use off the shelf ones) and we'd love to add an entire job graph to this queue and be able to query the state of it.
There is no way to define the dependency automatically, the workers can create any type of workflow though, including if Job X fails.
One of the things I awlays wanted beanstalkd to have was an atomic move-tube command, so you could emulate a state machine using queues and tubes
Internally Workq, for simplicity, does not have any separates "tubes" however. Just a job pinned by its name.
It's not a difficult feature to implemented (in fact if I recall there is a pull request open for this feature in beanstalkd) and IMHO would open up a lot of interesting use cases.
Looked very good though
What are the plans for persistence? Persist to disk? Or pluggable storage backhends? Disk, Redis, and SQL options would be cool!
There will be some sort of interface for the storage. I'll keep in mind pluggability. It is something Gearman had back in the day also[0]. Most likely it will be persistence to disk for some time and once more clarity comes out, possibly pluggability.
Workq is built on the higher level concept of a job so the feature set is refined around what a job is. In Workq, a job must successfully complete or fail, and optionally a result passed back to the client. A job can be retried when it has timed out or even when it has explicitly failed outright (maybe there was an temporary error with your API provider..etc). You can say: retry the job if it has timed out up to 5x BUT let it explicitly fail only once. These small refinements help streamlined the concept of processing a job fully, not just a blob of message.
Also there are some other key things such as job scheduling based on time which don't exist in RabbitMQ, but are usually offered in libraries such as DelayedJob..etc.
From my own experience, the text commands are easier to test against, especially its boundaries (inputs to the server) which makes client development significantly easier.
As for HTTP2, isn't that handled by the HTTP server/client implementations and is from the application code the same as HTTP?
As for the HTTP2 portion, the server details would be abstracted out especially with HTTP2 support in Go 1.6. At the time I looked, HTTP2 clients for various languages were still popping up and stabilizing (I think they still are). I didn't want that to be a factor when I was developing clients outside of Go (for example PHP). In addition, an important goal was to develop extremely small clients, where I understood exactly what was going over the wire.
If there are enough direct tooling benefits that HTTP2 can offer, it would be fun to experiment with it as an alternative interface. Funny enough, the first prototype name of the project was "httpq".
Thats not to say a HTTP interface is that much more difficult to add these days.
> Job payload & results are limited to 1 MiB each.
> Workq servers are standalone and do not speak to each other.
i.e. don't use it in production. This is a nice proof of concept, but let's not pretend that it is a professional grade product at the moment.
And nobody's ever going to trust your distributed capability if you implement RAFT yourself. You need to build upon something trusted or convince Aphyr to run Jepsen on your implementation. etcd is a common thing to build upon, and even it is only partially trusted.
I guess that could be said for a number of things that people successfully do?
Edit: for one great example of someone who didn't get discouraged by naysayers, check caddyserver.
And caddyserver isn't a distributed service, so the level of trust required is much lower.
Distributed services are difficult to get right for a wide variety of reasons, as shown by Aphyr's Jepsen tests.
Also I think the README of workq didn't mention anything about raft, which is fine, you can go very far without it.
My point is: encourage people to write code! Don't infect people with paralysis-by-analysis.
I spent 80% of time on Workq just writing tests for it and there is still so much to account for even as a standalone system.
On the bright side, projects like etcd & consul (I believe the author of Raft is helping out there) are getting better and better and can be embedded.
As they say, scaling is a nice problem to have.
We have been using beanstalkd for years to process gazillions of jobs without an issue and I guess many readers are in the same position.
Why would I choose Workq over beanstalkd ?
The one feature which may not be obvious yet is the ability for workers to mark a job successfully completed or failed with a result and then retrieve it later. The workflow looks like this:
* Client A: Backgrounds a Job A
* Client A: Backgrounds a Job B
* Client A: Backgrounds a Job C
----
* Worker A: Picks up Job A + Completes
* Worker B: Picks up Job B + Completes
* Worker C: Picks up Job C + Completes
----
* Client A: Picks up the result for Job A,B,C.
This allows a single client to concurrently process multiple jobs within a single process and retrieve its result. This is what I like to call "Gearman mode"[0], as it was modeled after that project also. Useful in languages that do not have well defined concurrency. This is a niche use case and may not be needed by everyone, but very useful as soon as you need it. This will become more obvious when I have clients for these languages.
Lastly there are some subtle enhancements such as retry support and synchronous processing (submit and wait for result).
Thanks for the great question. This is a very popular question and I will FAQ it.
For example: 68% done
Or: ETA: 15 minutes