I'm quite new to Ruby/Rails though — would be interesting to hear from others how they think it stacks up (unfortunately the author doesn't compare in the blog post)
I'm quite new to Ruby/Rails though — would be interesting to hear from others how they think it stacks up (unfortunately the author doesn't compare in the blog post)
Meanwhile, 250GB Postgres server has been churning along no problem.
Also, conceptually, architecture where state = DB, and business logic = web server is generally easier to reason about during service migrations and such.
If you don't use ActionCable which I think is the other core feature that relies on Redis, being able to remove it would've saved me a lot of early morning firefighting.
I've seen this pattern happen with a lot of Ruby projects: there's a popular Gem that people use that grows over time, then someone writes a blog post about why that package isn't suitable for their needs due to a design trade-off (often introducing a new Gem that is advertised as superior to the previous one) and then suddenly the old Gem stagnates, causing the maintainer to lose interest and updates stop pacing Rails versions. Sometimes the old Gem gets deprecated altogether and maintainers stick a big DO NOT USE sign at the top of the github README, all but guaranteeing there is no further community traction or organization for the project.
Meanwhile those of us using the old project chug along with a forked Gem that we cobble a bunch of patches into to meet our needs because there's no longer a centralized place to contribute to anymore. In some ways it's the double-edged sword of OSS. Maintainers aren't obligated to stay as maintainers of course, but it does get frustrating how quickly the community is willing to drop support of things that a lot people are using in production systems.
It definitely makes me reconsider the "don't reinvent the wheel" advice that is so adamantly thought of as a common sense convention in our craft.
There are very few “feature gems” that I’ve used in the past 15 years doing Rails development that I’ve been happy I used a year later. Development gems on the other hand like byebug are usually fine.
The problem here is the definition of moderate scale and huge amount of usage.
We handle tens of millions of events per year using our primary RDMS as a queue with absolutely no problems at all. The vast majority of projects for most companies aren't going to get bigger than that.
If you have millions of users sending out tens or hundreds of messages per day, sure, that's a huge amount of usage.
But I can see GoodJob being better for spinning up new projects, as there are fewer moving parts to worry about, and it's easy to upgrade to Sidekiq, etc. when you need to (thanks to ActiveJob's standard abstraction).
The creator and former maintainer of Redis was up until a few years ago discouraging its use as a queue, I think mainly because of its lack of durability and high availability at the time. He built a prototype Disque[0] to address the issues but it never became production ready. The other downside is that Redis is in-memory which means the queues have less capacity/are more expensive for the same capacity than an on-disk solution, but as memory gets cheaper over the years this becomes less and less of an issue. The upside is the throughput of Redis is very high.
I have personally worked on rails apps using redis-based queues like resque (and to a lesser extent sidekiq), and actually haven't run into any redis crashes or downtime in years of runtime, redis is very solid in general. You can also snapshot the redis instance periodically, to limit the number of jobs you would lose if it did crash.
In terms of using a primary db like postgres or mysql as a queue, I have personally run into issues with this multiple times. I would recommend never to do it, except on the smallest of side projects.
The issue is that eventually your queues will back up, whether it's due to a bug, surge of traffic, or just complex interaction of behavior in your app that cascades a ton of jobs at once when you run a backfill or something. When your app starts to get overloaded it's pretty trivial to increase the number of web instances running, so your bottleneck in these situations is going to be the db performance. As your queues get backed up, your queue workers are running at full speed processing jobs nonstop, which puts strain on your DB. Additionally, the act of enqueueing and dequeuing a job itself also puts strain on the db, so you can easily get into an unstable situation where each job that gets added to the queue makes every other job take longer.
If you allocate a separate DB instance that is only running your queue, that is much safer. Still, a DB like postgres is not great at doing constant writes and deletes, it creates additional auto-vacuum pressure for instance. But this will manifest as just getting worse throughput on the same hardware than you could get from a dedicated queue like rabbit mq, so if you're not at large scale it's a fine option.
Edit: And one other thing to add, for a lot of web apps the scope of what is needed from a queue these days is a lot less now than it was in the past. It used to be, and in large enterprise systems it often still is, the case that when people talked about a message queue they wanted something to facilitate passing messages between many completely separate apps. Now most apps just use a rest api for that (or perhaps protobufs or graphql or something but still over http). So I think historically an additional reason against using a simple datastore as a queue was that it didn't have enough features so you'd end up re-inventing the wheel with things like brokers, fan out and broadcast patterns, at-most-once vs at-least-once semantics, etc. But here I'm just considering the very limited usecase of a sidekiq-like queue, for processing jobs in the background for a single web app.
tl;dr: Never use your primary DB as a queue. Using a separate Postgres instance can work if you over-provision capacity and don't need to maximize throughput, and a Redis-based solution can work if you don't need high availability and can tolerate some messages lost if something goes wrong.
That is a broad generalization that assumes most applications are operating at mega scale. The benefits of simplified dependencies (a single database instance), transactional guarantees (a single database instance) and persistence (not using Redis) far outweigh the eventual possibility that the queue will place too large a load on your database.
As the author of Oban[0] (an PG backed persistent queue in Elixir) I'm definitely biased. However, the level of adoption in the Elixir community seems to signal that a lot of companies favor simplicity and safety over a possible scale issue down the road. The primary application I work on processes ~500k-1m jobs a day and the queue overhead is virtually invisible.
But in a past company I worked at, the company started out thinking sure just throw the queue in the primary db for simplicity. Eventually our slow query logs and db performance monitoring tools were showing that ~40% of the db load was due to the queue inserts and queries. It may be that it was doing something incredibly inefficient and unnecessary in the particular library we were using but we did look into it pretty thoroughly, this was a few years ago though so I don't recall all of the details. And that was at normal operation, then we ran into an issue that basically brought our site down when queues backed up.
At that point it was definitely time to split out the queue, and when we did it we realized that we had implicitly been depending on transactional consistency between the queue and the app data in a few places, which was then extra work to track down and fix these types of issues. This is IMO a code smell as well in general - your data and your infrastructure should ideally not be so tightly coupled.
Managed databases are so easy to set up these days, I would still definitely recommend a separate instance for the queue vs the primary db from day 1 in any new app I build. If you do want to combine them to save money on infra, use a separate logical db and separate connection pool and everything so that it's easy to split out in the future.
I will dig on your elixir queue, just build my own some days ago (mostly for fun) but also for solving some limitations in rabbit (mostly time based scheduling at short periods of time.. and control the throughput (to solve some rate limits).
Mine is here, https://github.com/vinissimus/jobs
The main feature is that it's built with pl/pgsql and allows to integrate so well with the rest of server backend (publishing jobs from triggers... ) Also listening on results with pg_notify
Sure there’s a point where this doesn’t make sense, but for the majority of cases there are ways to mitigate the positive reinforcement loop you’re taking about.
The easiest is to limit the number of workers to some number that won’t impact DB performance if they are running full tilt. You can even use an enum on existing records to determine the background job status if you have few enough rows (we do this for a table with a few hundred thousand job applicants).
I’ve found that in most cases we want a record that the background job was performed, so we were often updating the database anyway when a job was complete.
Sure if you are firing off so many events that PG can’t keep up with writing them, then PG isn’t a good option.
Postgres is fine as a queue for medium to large sized projects as it depends completely on your workload and what sort of performance characteristics you need. Thinking about it in terms of project size is the wrong way to evaluate it.
1. Ah, we’re using the database as a queue.
2. Database is running hot, I wonder where all the load is coming from?
3. Oh, it’s all the queue tables.
It’s not even a scale/load issue. Even at smaller scales you’re introducing some major lock contention.