Centrifuge: a reliable system for delivering billions of events per day
segment.com
segment.com
So if your webhook is bucketed into the 0-100ms queue and your responses start to exceed 100ms, you'd be bumped up to the 100-500ms queue which is more likely to have periodic queuing delays, and upwards from there depending on your response time / failure rate. If your API later recovered and started responding faster you'd be moved back up into the faster queue. That way we could offer different (soft) SLAs for different classes of response times, and scale the workers independently per queue.
I'm curious if there are known issues that people have run into with this approach? The main unknown was how many workers we'd be willing to throw at consistently slow endpoints to try to keep the slower queues from backing up too much, and possibly some flapping as endpoints could respond quickly under low throughput but slow down as soon as they move up to the faster queues.
If you want to slap it on your website for basic analytics, you better not get any traffic to speak of, because they charge you about $.01 per monthly tracked user, anonymous or identified.
If you get 100k visitors per month (not even THAT much), it's $1,125/mo. Their pricing estimator stops at $2,375 / month for 225k monthly tracked users.
God forbid you put it on your site or in your mobile app and then anything you do goes viral. It'll bankrupt you.
Who on earth is this aimed at? Is this only suitable for tracking logged in users? Or only suitable for enterprise companies?
Better yet, what about what they're doing is so difficult and expensive, when many of the places you'd be sending these analytics (which will be handling the same number of events) are free or cheap at this level of use?
Segment has always seemed cool, but overpriced for web analytics by a factor of 20-50x.
OK, rant over.
But it fills a really valuable hole for anyone that spends a decent amount on paid media. Having conversion and tracking data funnel through Segment has two benefits:
1. Every ad network has it's own conversion tracking system, and Segment makes it a whole lot easier to ad in another one that marketing wants to test out 2. It enforces a semblance of consistency. Implementation can have a very large impact on analysis. Whether intentional or incidental, it's really easy to implement analytics code inconsistently. Then the business/marketing user is evaluating results between campaigns/platforms/agencies/etc as if it's apples to apples, and it turns out it was implemented like apples to oranges. Standardizing events with Segment and then forcing tracking events based on those neatly ensures a level of consistency between services that's hard to achieve otherwise.
#2 is the biggest problem I've ran into managing clients' web analytics. My goto solution is to leverage a tag manager, work on mapping out significant actions and creating triggers for those, and then making sure as much as possible is driven by those standardized triggers and anything one-off is scrutinized appropriately.
But that process is easier said than done, since the friction to create new triggers is low and just requires appropriate GTM access. With Segment, making a completely new event involves an actual development release and is harder for a business user or agency to push through nonchalantly.
$0.01/visitor is not insignificant but in that environment (which is not all that rare), something like Segment is invaluable.
Plus, don't assume cost scales linearly with the same variable as publicized pricing. Have never been privy to a Segment contract, but part of the "Call Us" component is to ensure hte unit economics of the product make sense to your business. MTU is the unit they chose for automated pricing tiers listed, but there are other variables like the number of sources/destinations, number of seats, etc are negotiable.
High volume of MTUs but just a handful of sources/destinations? Can probably get a contract structured with a very low MTU price but a negotiated max active destinations with a per-destination charge over that. Your MTU peaky? Can likely get those to roll over or have the MTU measured against a per-month average over a quarterly basis. Allowing your peaky periods to be offset by your down periods. Don't view the Business custom pricing option as only for super huge spenders - a lot of companies are willing to work with you when you're smaller in hopes to structure a contract that'll grow annually as you grow. Custom pricing doesn't always mean super expensive, it also means "atypical unit economics". And also happens to be a market research tool to help companies discover use cases they hadn't thought about (hence why their standard pricing differentiator didn't work for it).
Think of it as a central data store for your business that has built-in integrations to move that data to other providers such as Google Analytics, Salesforce, Mailchimp etc.
But if one of those services is Google Analytics, for example, then you probably want to use it with all visitors, which is excruciatingly expensive if you get hardly any traffic at all.
I've been posting this a lot recently but it keeps coming up as the relevant solution, Apache Pulsar supports millions of topics without much overhead and offers the log semantics of Kafka with better scaling and per-message acknowledgement: https://pulsar.incubator.apache.org
I would imagine you're taking the database version from Consul and using that as a fencing key for writes within a transaction in the database, but I didn't see that mentioned in the Go code snippet.
I fail to understand the following aspects though, maybe you can clarify them a bit:
* How does a director recover from a failure? From my understanding it would require fetching all job IDs from jobs that are in an active state via the job transactions table (which sounds expensive) and then loading the associated meta-data from the jobs table? Is that correct?
* Do you assume that the director will have archived all non-completed jobs when deleting the database? Do you try to gracefully shut down the director first then? From my understanding it seems you perform a "drop table" statement on a given database and then regenerate the tables, but this would require being sure that all the jobs have been processed or archived.
> How does a director recover from a failure? From my understanding it would require fetching all job IDs from jobs that are in an active state via the job transactions table (which sounds expensive) and then loading the associated meta-data from the jobs table? Is that correct?
This is correct, if a director crashes for any reason, it needs to scan the database on boot. It is a more expensive operation, which is part of the reason that we try and cap the number of total entries in a given database.
> Do you assume that the director will have archived all non-completed jobs when deleting the database? Do you try to gracefully shut down the director first then? From my understanding it seems you perform a "drop table" statement on a given database and then regenerate the tables, but this would require being sure that all the jobs have been processed or archived.
Great question, we glossed over this aspect this a bit in the post itself.
Before a given database is transitioned to the 'spare' state, and its tables dropped, a single Drainer process is responsible for moving any non-completed jobs from that database to another active Director. The Drainer will not successfully exit and transition the database to 'spare' until it is certain it has processed all the non-completed jobs. We never drop any tables which have non-terminal jobs. Similar to the Directors, the Drainer will acquire a lock in consul to ensure only a single process is draining at a time.
We're hoping to go into a bit more depth on how the drainers work and these jobs move around in an upcoming post on Centrifuge's two-phase commit semantics. Ensuring that your data has moved to another system does require fairly complex transactional semantics, so we're hoping to go into depth about how this works.
On your second point, archiving actually happens in the “drainers”, not the director. It’s the component that picks up unused databases and flush their data back into the set of running directors. We initially built archiving into the directors but it turned out to steal too much resources away from job processing so we moved it out into the drainers, which actually gives us opportunities to do it more efficiently (for example creating larger archives, which helps with compression as well).
Sorry if we cut some details off of the post, there is plenty more to tell about this system but it’s a lot for a single blog post ;)
I have been on the lookout for a similar system to SideKiq.
Unfortunately for Go, there isn't anything that matches it 100%.
I have looked into:
https://github.com/celrenheit/sandglass
https://github.com/RichardKnop/machinery
https://github.com/contribsys/faktory
In the end I went with machinery due to supporting:
- Batches (or Groups, tasks executed at the same time)
- Chains (tasks executed one by one)
- Callbacks (a task executed, on completion of a Batch)
- Cron Jobs
- Rescheduling failed jobs
- Long-runnng Jobs
- Distributed/Fault Tolerant DB (for Jobs)
- Distributed/Fault Tolerant Workers
+ other things I cant remember now.
It would be very interesting to know if this is going to be opened sourced and whether it supports the list above.
From their own ReadMe:
> Fairway isn't meant to be a robust system for processing queued messages/jobs. To more reliably process queued messages, we've integrated with Sidekiq.
It's a supervisor (Director) with a bunch of actors communicating with different 3rd party API's. The state mechanism could ostensibly get abstracted away so the underlying DB is irrelevant.
It has to deal with lots of third party interfaces, websockets, webhooks, etc. Erlang is notoriously bad for modern web stuff, with elixir making the ecosystem more brittle and lower quality. If you're building a system that needs to do a lot of back-and-forth, go is a much more usable tool.
Have you ever tried to debug a giant erlang clusterfuck, calling out to a dozen different service gateways, varying tuple msging nonsense all over the place, jankedy ass comment annotations that aren't even accurate half the time, and all the while trying to fix things with hot swaps?
It all sounds very good when Joe writes those weird little lewis carroll love letters about how magical erlang is, but I've worked on a system like that and I'll never touch erlang again except for small single-shot services. It's completely cancerous as a language ideology.
Not that go is all that much a panacea, but it's definitely the perfect fit for this sort of thing.
Hope this doesn't come across as too harsh, but what you're describing is a rabbit hellhole I've already crawled through.
Also, with projects this size, abstracting db details isn't wise. This kind of scale necessitates more precision in what schemas you choose and what configuration your dbms has. Crap like ecto will ultimately cost you in efficiency. Elixir can get away with it because beam works hard to do supervision thoroughly, but there are multiple performance penalties that beam imposes on itself by choosing abstraction instead of specialization, and again, this problem would demand a specialiZed approach to be efficient as it scales up.
> To keep the ‘small working set’ even smaller, we cycle these JobDBs roughly every 30 minutes. The manager cycles JobDBs when their target of filled percentage data is about to exceed available RAM.
I'm confused - why does JobDB's memory trickle up over time? Isn't it a database? Are you using MySQL's memory storage engine or something?
I hope this gives some clarity, let me know if you have further questions.
In a similar position I would have tried MQTT that should handle quite well a great number of topics.
Did you guys tried such protocol?
You build your infrastructure on top of it to support stuff like storage, retries, resiliency, etc...
However, it does gives you some guarantee, namely that the message will oblige its QoS (at most once, at least once, exactly once).
On top of these guarantees, you build whatever you need.
Honestly, I am quite glad that you are writing these posts, it is a great service for the community and if you end up open sourcing it, the project will bring a lot of value to everybody.
Thanks for your posts :)
https://github.com/centrifugal
It looks like not, but that's a hell of a naming overlap.
I looked at Centrifugo a few years ago to deliver live Hearthstone games (in game replay format) through the web. It's a pretty sweet project.
1: https://github.com/segmentio?utf8=%E2%9C%93&q=&type=&languag...
BTW, it also appears there is an active US trademark (#78010053) for "centrifuge" in the computer database context. Isn't that an issue, too?
I started doing that after having to go through two iterations of renaming a project after finding out the old name was somehow encumbered... of course after we had already used it in front of other people; it was very embarrassing.
Still, do keep in mind that it's not the name of a product offering - it's solely an internal name for the technology.
It started with "Cucumber" or "Celery" or something a few years ago didn't it?
I'll skip ranting about how "Go" was a terrible name for a PL (and only mention parenthetically how they gaffled that name from a different PL!) They have "Grumpy", and something called "Thanos" (good luck searching for that until the hype for that comic book movie dies down.) I feel like I've seen several other projects recently that have been named after other things.
And Elm, Swift, Logo, Pascal and Haskell and Ada, Icon and Self, Opal and Occam and Maple...
There's a lot of them: https://en.wikipedia.org/wiki/List_of_programming_languages
If all your friends name their projects after something and then jump off a bridge, it's still a dumb thing to do.