Procrastinate: PostgreSQL-Based Task Queue for Python
procrastinate.readthedocs.io
procrastinate.readthedocs.io
I'll leave my questions here in case people would like some food for thought anyway.
- what's with the mixture of styles (q{sql}, sql with '?', sql with '$1', sql with '\$1' etc)?
- what's the use case for a right join, when left joining is the norm?
- in the $stats query, inactive_workers isn't actually inactive_workers and instead it's patched based on active workers afterwards. For me, the jumping of inconsistent column selections on each line makes it hard to see that. (If someone just copied the sql into psql to run it, it would be wrong, but not obviously)
- Maybe it's not fair, because I have sqlalchemy to improve readability, but this whole thing where the variables aren't obvious in a blob of sql could definitely be easier to read
While the beauty is in the eye of beholder, and you may be correct, or have high standards, praise coming from the likes of @mst should be taken seriously. People may have seen code that can stand in for sh?t, so it makes them appreciate simple niceties like good formatting, good and concise logic, etc.
BTW, I have to agree with @mst on this one; even though perl is very low on my scale of pleasantness for languages, the SQL in that code is beautifully written.
I’ve upvoted you partly on those grounds, and partly to get you back to 1337 :)
Thank you!
So I guess that while you perhaps did have suboptimal judgement in your initial reply you very much weren't alone and I should remember to be more careful about that going forwards.
Anything specific worth calling out?
Certainly, very little of my own code passes that test after I've not looked at it for a week.
I thought it seemed like a project deserving a bit more attention.
I was using Celery for sending emails - nothing else. And Celery was such a nightmare to configure and debug and such overkill for email buffering that in a fit of frustration I wrote the Arnie SMTP buffering server and ditched Celery.
https://github.com/bootrino/arniesmtpbufferserver
It's only 100 lines of code:
https://github.com/bootrino/arniesmtpbufferserver/blob/maste...
And why Postgres and not rabbitmq or redis?
1. For anyone whose stack is currently a simple three-tier architecture that currently has Postgres and nothing else in the DB tier, adding anything else would incur 100+% operational complexification (e.g. now you have to figure out how to do backups for two stateful components.) Far more than 100%, even, if they get the Postgres DB "for free" as part of a virtual LAMP-stack appliance, or built into the base offering of some PaaS service, and currently don't need to do any ops work as the framework/platform handles their DB's care and feeding for them. (Ideally, such setups would do the same with a built-in MQ as well — but sadly, that's much rarer.)
2. You can be clever with a job queue that’s in your DB, by making “taking the job” and “doing the queries that comprise the job” part of the same atomic transaction, such that a ROLLBACK due to a constraint validation error in "the queries that comprise the job" will also implicitly “put back” the job immediately (even if the worker crashed in response to seeing the error.)
If you use Celery, the time saved dropping Celery will more than cover the operational costs of redis.
Really, one of the magic tools of the 21st century.
I chose postrges for our mq stack over rabbitmq and redis. Since our team is small, we are trying to be efficient in our use of manpower. our architecture is as dirt simple as it gets. (loadbalancer -> api-server-nodes -> aws aurora).
adding another external dependency just adds one more thing that can potentially go wrong. Thats one more thing we need to watch and maintain accross staging and production envrionemnts and one more thingto get running on developer machines.
There may come a day where we switch over to rabbitmq or kafka but postgres has proven to be fast enough. when writes are too much of a load, we can shard the writes to its own dedicated machine and buy more time. When our traffic is sufficiently high that even that isn't enough, we'll have enough revenue to pay someone to configure kafka/rabbitmq and deal with the configuration and maintenance fulltime.
Until then, if you want to survive as a startup, KEEP IT SIMPLE. Any complexity you adopt needs to justify itself by providing a competitive advantage.
I love the idea of re-using infrastructure "for free".
I have a production system running on Celery (on Redis). It's been fine, once I tuned it.
Time for a refactor; wondering: should I consider migration?
In the past I have actually written my own very little scheduler that borrows the app database. Really just table that is a job queue and a script on a loop that grabs the job when it appears in the table. But it did have the occasional hiccup.
I can't wait to check this out and some other suggestions in the comments.
It's worth mentionning Huey too: https://huey.readthedocs.io/en/latest/ I've been using it for a year on top of sqlite.
It's not open source though it is free.
StarQueue has a pure HTTP API so any language can talk to it.
StarQueue is designed to be a more simple implementation of the SQS way of doing things, without some of the unnecessary stuff that SQS has in its API.
Under the hood it uses Postgres although I actually wrote it to support MySQL and SQL Server also, which are equally capable of doing queueing as Postgres.
Is it: - They want persistence in case the system goes down, or the task queue service is restarted.
- They want to replicate their task queue service and want to be sure that the data is properly synchronized.
- The amount of tasks/time if simply too much for some standard library queues
- They think is is easier to add another dependency than implement it themselves
- something else
?
It sounds like your question is focused on "why not just use an in memory queue". Two reasons: 1. You loose messages/jobs, and it's not scalable 2. If the task is CPU intensive, it will lock up or slow down your web server.
And when it (the database) cannot serve your needs in terms of performance anymore, then you must choose a domain-specific product dedicated to solve the you challenges.
> premature optimization is the root of all evil.