Faktory, a new background job system
mikeperham.com
mikeperham.com
I guess this style significantly reduces setup friction, but it’s almost an irresponsible design in a universe where the average cloud provider is telling you, “your compute instances may vanish at any time.” If this used Kafka, Redis, MySQL, or another well-known stand-alone data store, I know my data is replicated, and I can recover my enqueued jobs if Somethimg Bad happens.
I like the nice UI and the simple API, but there’s no way this will replace Resque or any Kafka-based-job-whatever at my place of work.
Side note: Resque will lose jobs if it crashes, it doesn't use RPOPLPUSH. I hope you'll take another look at Faktory in a few months, maybe we'll have addressed your issues.
There are three common solutions to this:
First is to pretend it doesn't exist (quite common!).
Second is to understand that losing messages is a possibility and only use it for message types that the loss of a message wouldn't be critical (which covers quite a bit).
The third solution, which actually addresses the problem, separates the persistence of message details to transactional store from the event notification of a new message. A "sweeper" type task can thing check the persisted message list for messages that haven't been processed and re-publish them.
(Disclaimer: I'm a huge fan of Redis and think it's the bee's knees of data structure servers.)
It's possible to fsync on every write[1] but it may be too slow as the single threaded nature of Redis means you've serialized every write operation. Plus that'd apply to all usage of that Redis server, not just the queue.
[1]: https://redis.io/topics/persistence#how-durable-is-the-appen...
What exactly is the use-case for an entirely new and separate system that looks like it's single-node only?
If your current system is already something like SNS/SQS or an actual message queue with acks, then this probably isn't aimed at you.
Push, fetch, ack = insert, select, delete. A few lines of SQL or a stored procedure gets it done. Postgres basically made skip locked for queues.
Also worth noting, postgres is also another piece of infrastructure to run, and not all applications use it.
For postgres, SKIP LOCKED is the easiest way to queue because the row only gets deleted if the transaction commits, which is when the job finishes. Otherwise if there's an error than don't commit or just crash and the row remains in the queue and another worker process tries it again.
If pg works for your situation, that's great, but i don't know why you'd assume that none of these objections have come up before and been responded to.
They couldve just ported the sidekiq library to several languages or use a single C library with language wrappers and gotten the same result with less work.
There’s definitely a lot of message queues out there these days and things like Kafka which turn into message queues. That said, a lot of them are a pain to operate so there’s room for one that is easy to deploy and operate.
This breaks the devops story for me. The difference in featureset between Rocksdb and redis is not that big.. However redis is hugely supported on the cloud and in high-availability fully-managed mode.
Its so convenient to use sidekiq on heroku or aws. I really hope you build this on redis rather than a new persistence server.
From my experience, when running high throughput, quickly executing sidekiq workers on heroku, the expense often doesn't come from dynos, but the redis instance as the limiting scale factor usually comes down to connections. That won't be a problem with Faktory and an embedded datastore.
Which means that leveldb/rocksdb gives zero path to scalability at the cost of saving a maximum of a few minutes of effort.
At worst, this could have been done in redis with embedded lua packaged together. That would have solved both your external redis problem as well as a reasonable path to scale.
> The Ruby and Crystal versions of Sidekiq must remain data compatible in Redis. Both versions should be able to create and process jobs from each other. Their APIs are not and should not be identical but rather idiomatic to their respective languages.
It makes sense to unify the protocol for background job processing, making them language agnostic. Some languages tackle different problems better than others, so this will be a really useful tool.
Great work, Mike. Keep the hits rolling.
Maybe for something like RabbitMQ or SQS, this would be a satisfactory replacement, since these seem to be relatively monotasked persistence servers. So for your traditional Celery+RabbitMQ deployment, for instance, this could be a good replacement.
But let's we consider cases like Redis, Memcached, or Kafka, where the persistence store we're using is often also being utilized as a cache or linear log in other aspects of the same product. This would make Faktory troublesome, because it introduces additional maintenance costs compared to a service that we already need. Furthermore, if I can use an off-the-shelf, hosted storage solution for enqueued job definitions, like we do with Elasticache, I reduce my operational costs even more.
So it's not a question of whether or not Faktory works, but whether it's worth the cost of building, deploying, and maintaining a specialized monotasker server instance on top of the worker pool I already need to build and maintain. I'd be interested in understanding where the long-term value add would be in most large-scale practical SOAs, and how folks excited about this project anticipate implementing it might go.
Although to be honest, the basics of a work queue are well covered now with the evolution of cloud services and other databases and message systems.
There's many different tools and many different users. No one choice is appropriate for all. I hope some people find Faktory useful.
Your potential customers know better than to get their opinions from reading these comments :)
I really respect what Mike’s done being able to monetize his work on cool open source projects.
The other thing about job systems that a lot of people seem to ignore is client-side scaling. We run our apps on Kubernetes, where you'd naturally want to tune worker scheduling dynamically to accomodate queue size. Feeding custom queue metrics into the horizontal pod autoscaler is one way to do this.
What can a job queue do that Kafka or RabbitMQ cannot do?
There are some useful features like priorities, scheduling in the future, claiming and then acknowledging a job as done or returning to pool, etc... but these can all be found in or coded on top of existing systems.
> Faktory aims to be more feature-rich and better supported. Many of Faktory's OSS competitors are "dead" and no longer supported. I am fortunate enough to have both expertise in background jobs and a business model to support Faktory long-term.
Not saying that Gearman is a better solution, just saying that "better supported" for such a relatively simple tool is not always necessary.
My question is why to we even need a 'background job' system, isn't this just a message broker / queue? RabbitMQ (and friends) can do much of this, no? Maybe I'm missing some of the future features they aim to implement.