BlazingMQ: High-performance open source message queuing system
bloomberg.github.io
bloomberg.github.io
I worked on this team very briefly. Great team working on some interesting tech.
grpc client a -> message broker(implemen a) -->
grpc client b -> message broker(implemen b) -->
message broker-> your app for the broker this make sense if dont make sense manage the quque in your app if ypu want to abstract the implementations is worth if you have n to 1 or n to n and all works different
If you need light-weight reporting like a status code you can use RPC/reply mechanisms if the broker supports it (e.g. rabbitmq), build your own (custom headers), or use an API/gRPC call.
Everything else starts needing some sort of more sophisticated coordination mechanism and gRPC will look attractive for that although there are usually other options as well.
This leads to a whole heap of benefits. I've outlined a few of them here, in the "Rationale for Mats3", my Message-Oriented Async RPC library: https://mats3.io/background/rationale-for-mats/
You will find that the Reactive Manifesto also lays out the same type of arguments (there's some links to it from the list in the first link): https://www.reactivemanifesto.org/ and its glossary: https://www.reactivemanifesto.org/glossary
This is very interesting. One of the benefits of Kafka is partition routing so I'm curious to learn how this may compare.
Cloud Providers can be expensive.
Do you have a story around this? If so, would love to hear more about it.
Article "AWS’s Egregious Egress" (2021) https://blog.cloudflare.com/aws-egregious-egress/
HN discussion: https://news.ycombinator.com/item?id=27930151
Multicast puts the entire load of distributing the data in its natural place: the network. The pub/sub queues that the rest of the world uses are more complicated and a lot worse.
This appears to be the beginning of a "cloud-native" bloomberg that can't lean on multicast.
Message queuing is a more complicated system, but it allows you to push the bulk of your work out to the distributed network instead of concentrating it on the producer.
You have to do this state management with a TCP-based system anyway - the use of TCP doesn't magically come with the full state of past messages.
By having a proxy/repeater/site server, you are reducing the amount of WAN traffic you need to send.
low cost end replicas at the end. if the bandwidth usage is half the cost in their multihop topology, 14XXXX compared to 16XXXX? depends on the number of proxies or replication and pricing.
Has anyone used comdb2? It seems like an awesome DB that doesn't seem to get the level of love it should.
BTW, the original Comdb2 paper from 2016 is good [0]. It has features that even today seem awesome. It sounds like a really nice distributed MySQL/PostgreSQL.
Also, it's worth identifying and setting aside use-cases for broker-less messaging where 0mq, nanomsg, etc. could be more appropriate.
Edit: There's only published performance data of itself, which has little value: https://bloomberg.github.io/blazingmq/docs/performance/bench...
BMQ license is apache 2.0 for anyone curious
I.e. selection of ideal server count is a multi-dimensional problem.
Yes, 4 servers might handle more load than 3 servers [1], but so would 5 servers. I'd guess that if you've got enough message queue traffic that 3 servers can't handle the load, those servers aren't a significant fraction of your total cost. You might as well go to 5. Likewise, if 5 aren't enough, I'd skip over 6 and go to 7. If 7 aren't enough...I'd probably rethink my design entirely...but I'd prefer 9 to 8.
Alternatively, quorum systems can support "zero-weight replicas" which effectively don't vote but still can be useful for eventually consistent reads. Then you have this server that helps handle the load but doesn't increase the number of servers that have to participate in a round.
[1] if the load is overwhelmingly eventually consistent read operations and their multi-hop topology with proxies doesn't sufficiently reduce the load on the replicas. It's not obvious to me if this scenario is plausible or not.
As seen in their source code here: view-source:https://bloomberg.github.io/blazingmq/assets/animations/anim...
1. Build the visual content in an SVG with Inkscape. It should be easy for you to build the layout with assets/item to animate. Don't forget to group items and set names so you can reference them later.
2. Once exported, insert the SVG in an HTML page and add anime.js. This library will let you animate the content you just created (you should be able to reference your named items from Inkscape and animate them). The learning curve might be tough depending on your experience regarding animation (or CSS in general).
Good luck :)
While I like the idea of something like anime.js and think it raises the ceiling on what’s possible, I’d still probably start with something like keynote just to get the idea out of my head and into a format I can validate before going into code.
Pretty sure it was made by Airbnb, and used in tons of apps like Google home
Works really well
The current standard is NATS JetStream, which absolutely addresses scalability, arguably to an even larger degree than Kafka, since it can be used to create a globally interconnected supercluster of clusters where messages can be configured to flow between regions as needed. I doubt anyone would run a Kafka cluster that spans multiple regions.
Stretch clusters and clusters with brokers in different AZs are absolutely a thing in Kafka. I make no comment on the ease relative to Nats, which I don't know at all.
EDIT:
Definitely sounds like there are some tight latency requirements[0]:
> A stretched 3-data center cluster architecture involves three data centers that are connected by a low latency (sub-100ms) and stable (very tight p99s) network, usually a “dark fiber” network that is owned or leased privately by the company.
100ms would theoretically be enough to span the contiguous United States, but the references to "very tight p99s" and "dark fiber" make me wonder if 100ms is actually acceptable, or just a theoretical maximum that can only be allowed under absolutely perfect conditions. Either way, suggesting the user should have access to dark fiber between their regions does not fill me with confidence about the robustness of this solution. I'm sure it would be fine for geographically close datacenters, as AZs are designed to be.
[0]: https://docs.confluent.io/platform/current/multi-dc-deployme...
Mostly because of the speed of light, not because Kafka is for some reason unable to form a cluster across regions but Nats can.
My main gripe is lack of “best practices”. You’re pretty much on your own when it comes to important decisions like subject namespacing and some operational stuff can be a bit tricky, like changing options and migrating. They had “nats by example” but that seems to be quite out of date and not super helpful. They have started doing more videos which are great but not that much content overall.
Their auth story is very sophisticated but perhaps a bit over-engineered. Some use cases like having anonymous “self-signed” client keys are tricky to work with.
Other than that, it’s an amazing piece of tech. One that is scalable in a way that just makes sense and gets out of the way when multiple machines are involved.
They're fundamentally different designs.
Garden variety NATS (i.e. the non-"Jetstream" case, hereafter) is "at most once" (messages will be lost.) BlazingMQ is "at least once" (>1 delivery of a message will occur.)
That difference is down persistence. NATS doesn't persist messages to stable storage. BlazingMQ must use persistent storage. From that all sorts of application design, performance and other implications are derived. And they all matter. A lot.
"Every byte counts" was the motto of the team that designed it.
I know that's not an answer but if you're feeling adventurous you might be able to help the project.
I'd tolerate wrappers for certain things, but not for this thing.
Its current sole implementation is based on Java Messaging Service JMS API, and it is used in production for a rather large UCITS unit holder registry on Apace ActiveMQ, and all tests runs fine on Apache Artemis MQ.
Every time a new message broker comes along, I sit up in the chair and wonder a) about their performance (!), and b) whether they have a JMS client implementation, and c) whether Mats3 works with that! When I tested RabbitMQ's JMS client library, I sadly found that there was rather many differences - things I thought was screamingly obvious, was not available. E.g. as basic function as redelivery: "Normal" MQs try to redeliver N times, typically with a backoff between each attempt, and then, when all N attempts fail, it puts the message on the DLQ. Rabbit instead tight-loops the delivery attempts until either the message goes through, or the sun burns out. To be able to use Rabbit, I would have to use the native API, and implement redelivery and DLQing myself, client side. Also, transactions.
I now wonder whether I should make a lower-level abstraction, so that the JMS implementation is converted to a "base" impl, and then the JMS, Rabbit, Bloomberg, NATS, ZeroMQ, Aeron, etc implementations was extensions, or "plugins", or "drivers", to that.
I'm also having trouble figuring out if Mats3 is a library (with a JMS API) over a variety of messaging systems (WebSockets, NATS, etc.)?
P.S. Some diagrams like https://bloomberg.github.io/blazingmq/ would be very helpful, especially at https://mats3.io/background/what-is-mats/. If a picture's worth a thousand words, and an animation must be worth at least 10k words. :)
There is an illustration on the front-page: https://mats3.io/
This page tries to directly explain the idea - but I guess it assumes that the reader is already extremely familiar with message queuing? https://mats3.io/docs/message-oriented-rpc/
Here's a set of small answers, from different angles, to "What is Mats?": https://github.com/centiservice/mats3/blob/main/README.md#wh...
Here's a way to code up Mats3 endpoints using JBang and a small toolkit which makes it extremely simple to explore the ideas - the article should also be skimmable even without a command line available: https://mats3.io/explore/jbang-mats/
If you read these and then get it, I would be extremely happy if you gave me a sentence or paragraph that would have led you to understanding faster!
The concept of messaging with queues and topics are essential to Mats3 - but the use of JMS is really just a transport. I could really have used anything, incl. any MQ over any protocol, or ZeroMQ, or plain TCP - or just a shared table in a database (but would then have had to implement the queue and topic semantics). As a developer, you code Mats3 Endpoints by using the Mats3 API, and you initiate Mats3 Flows using the API. You need an implementation of the MatsFactory to do this, and the sole existing is the JmsMatsFactory - which works on top of ActiveMQ and Artemis's JMS clients.
Wrt. WebSockets, that is a transport typically between a server, and a end-user client, e.g. an iOS App. Actually, there's also a "sister project", MatsSockets, that bring the utter async-ness of Mats3 all the way out to the client, e.g. a webpage or an app. https://matssocket.io/
NATS is, AFAIU, just a message broker, with some ability to orchestrate. Fundamentally, I do not like this concept of orchestration as an external service - this is one of the founding ideas of Mats3: Do the orchestration within each service, as you would do if you employed REST as the ISC mechanism. I do however assume that one could make a Mats3 implementation on top of NATS client libs.
#) Java's Project Loom is extremely interesting to me, as its argument for using threads as the basis of concurrency instead of asynchronous styles of programming, is exactly the same rationale for which I made Mats3: It is much simpler for the brain to grok a linear, sequential, "straight-down" code, than lots of pieces of code that are strung together in a callback-hell. Async/await is a way to make such callback-hell more readable, "linearaize it" (but you still have the "colored methods" problem which Loom just obliterates in a perfect explosion). One could argue that this is what I have achieved with Mats3 when using messaging.
C++ and Java client examples are very simple which is cool.
Would appreciate if there is hook to enable `SO_TIMESTAMPNS` on socket and retrieve `cmsg` to get accurate timestamping and permit client code benchmarking.
> In a zero-replication no-leader-election (or all-in-one-node) scenario what is min latency observed?
We saw sub-millisecond latency in this case (typically double-digit microseconds).
Kafka has its uses, in particular for massive influx of events, e.g. in a large IoT system - I'd say the perfect example would be continuous measurements of tens of thousands of gauges and sensors on an oil rig or any other large production system, or e.g. the energy meters in every home.
But I would personally never use such a system as the inter-service communication layer for a multi-service architecture. It is WAY to heavy coupling. Event sourcing looks fantastic on the surface, but is a disaster for a decades-living system with tons of developers. REST/gRPC is better, but async messaging really rocks.
This goes smack in the face of the idea of "one service, one database" 1), where there is a distinct boundary where each service owns their own data. How a given service stores it data is of no concerns to any other service - the communication between them is done using e.g. REST or messages, with a clear contract.
Event sourcing / Kafka architectures are the exact opposite of a clear contract. You are exposing the absolute, deep-down innards of the storage solution of a service, by putting its microscopic events down for all to see, and all to consume. You may do aggregations, and emitting more coarse-grained "events" or state changes, thus kinda also exposing "views" of these inner tables, and maybe use a naming scheme to sort of which are "public" and which are not.
In the beginning, I really did find the concept of event sourcing to be extremely intriguing: Both the "you can get back to any point in history by replaying up-to", and "forget the databases, just emit your state changes as events" (I truly "hate" databases, as they are the embodiment of state, which is the single one thing that makes my field hard!), the ability to throw up a new projection for anything you'd need (a new "internal view") of the state, and that unit testing could be really smooth.
I upon deeper delving into the concepts found that the totality of such a system must quite quickly become extremely heavy: The amount of events, thus needing snapshots. Evolution of events, where you might have no idea of who consumes them (that is, the "shared database" problem). The necessary understanding, throwing a half-century or more of relational databases under the bus. The performance tweaking if things start to lag. Etc etc etc. It would become a distributed system in the worst way a distributed system could be distributed: All state laid out in minute details, little abstractions, and a massive diverse set of different implementations in the different services to get back to a relational view of the data. And this is even before the system gets a decade old, with lots of different developers coming and going, all of them needing to get up to speed, and the total system needing extremely good and hard steering to not devolve into absolute utter chaos.
That Kafka can be employed and viewed as a database has been argued hard by Confluent. Here's their former DevEx leader Tim Berglund explaining how databases are like onions, and that in the base of every database there is a commit log. And that this log is equivalent to a set of Kafka topics. 2) So why not just implement all the other layers of the database in application logic? Confluent even have a lot of libraries and whatnots to enable such architectures.
1) "Microservice Architectures", patterns: https://microservices.io/patterns/data/database-per-service....
2) JavaZone Talk from 2018 by Tim@Confluent: https://2018.javazone.no/program/3a9644e6-15b5-4c66-a28c-c35...
I imagine this is used for Bloomberg (the terminal) and not Bloomberg (the website)?
going back to the article - fantastic animations. I'm just as curious to how that was made as the queue itself.
Of course they do, unless the processes are regulated by a country (such as handling highly addictive opiates in the US).
Then there is language variances like flavors so what ends up happening is a million different conceptual primitives that fundamentally the same.
It isn't the same because of a few things. At the end of the day too it is also about how good the support is over the rest. If a new MQ can demonstrate better tooling nothing stops a company from adopting it.
If anything, we should be thankful that they've decided to share it with the world.
It's a big driver of both complex and slow software. Often the adaptation layer necessary to get it to actually satisfy your needs rivals the size and complexity of the piece of existing solution code you're trying to leverage.
Software that is built from the ground-up to your exact needs is often both faster and slimmer and more maintainable and suffers from far less dependency churn.
In the original sense, the NIH approach would be to invent a rivalling technique to the message broker. That is indeed a questionable idea.
Seen Not Implemented Here too.
Well, no one cares, but I'm not saying that to be mean to dismissive, it's just that software production as an industry (esp. at a company like Bloomberg) can basically do whatever it wants because of low capital costs to produce something. Things like writing another queue system will have zero impact on the industry unless it's adopted widely (and cargo culted) and later made worse.
Maybe I'm about to eat my shoe, but if a capable of team at Bloomberg wanted to reinvent something considered unnecessary in terms of industry need, they probably could if for no other reason that management allowing some indulgence in exchange for keeping talented people around is Worth It™.
They do adopt a lot of industry stuff, but it is always heavily modified. And of course they love “Invented Here” as a corporate tradition going all the way back to the beginning.
It’s actually nice from a developer perspective, the R&D group is highly valued by management.
I don't understand that phrase. Probably a large team of very well paid engineers developed this product. Bloomberg is well-known in the industry for paying well. "[L]ow capital costs"? This software probably cost millions in salaries to build, and millions more to maintain over the next 10+ years.
>I don't understand that phrase.
You don't need a big factory with large capital costs to make or distribute software. You're only paying people, and salaries are not capital investments.
1. It is easier to implement from scratch than understand and modify: we see this everyday at a way smaller scale with libraries/classes/functions. Not saying this is optimal but this is a driver for sure.
2. Recruiting: some devs love doing this. Fancy OSS project to put on your resume, some actually believe there's unfulfilled need, you get people like us discussing it (otherwise we wouldn't be thinking about Bloomberg).
I've polled developers about what part of software development they hate, and in general, it's deployment and maintenance (ie: devops). Using a off-the-shelf-component means that the developer only gets to do the un-fun part: installing, configuring, productionizing. If they code their own solution, they get to spend a few weeks coding!
I am wondering if they ran into performance or resource usage issues with Kafka at scale, and decided to roll their own for use cases that better fit their workloads.
Bloomberg does have vast backend infrastructure for moving around / computing data which the Terminal is a consumer of.
If software is written 100% externally, but the vendor only wrote 10% of it and is transparently taking 90% from an open source project so they can sell fraudulent support (assuming they don't have people contributing to the open source project here), that counts as supported.
If it is written 100% internally, even if it is built on top of some ancient internal codebase that nobody can figure out and has been in life support mode for 20 years, then it counts as supported.
But if its written 50% internally and 50% of it came from an external open source library with a BSD license, and you are a major contributor to that project, then it is considered unsupported and you get in trouble.
This is how you end up with security issues because some sysadmin decided they had to roll-their-own crypto library or authentication system, because its better to have an unknown implementation show up on a scanner than a known one that has a list of CVEs that can be checked.
https://devboard.gitsense.com/bloomberg
For the 200 repos that they have made open source, they had over 6000 people contribute feedback and/or code. So in addition to attracting talent, they also get more testers and feedback.
Full Disclosure: This is my tool
So their own homegrown leader election algorithm?
> BlazingMQ’s leader election and state machine replication differs from that of Raft in one way: in Rafts leader election, only the node having the most up-to-date log can become the leader. If a follower receives an election proposal from a node with stale view of the log, it will not support it. This ensures that the elected leader has up-to-date messages in the replicated stream, and simply needs to sync up any followers which are not up to date. A good thing about this choice is that messages always flow from leader to follower nodes.
> BlazingMQ’s elector implementation relaxes this requirement. Any node in the cluster can become a leader, irrespective of its position in the log. This adds additional complexity in that a new leader needs to synchronize its state with the followers and that a follower node may need to send messages to the new leader if the latter is not up to date. However, this deviation from Raft and the custom synchronization protocol comes in handy because it allows BlazingMQ to avoid flushing (fsync) every message to disk. Readers familiar with Apache Kafka internals will see similarities between the two systems here.
"a new leader needs to synchronize its state with the followers and that a follower node may need to send messages to the new leader if the latter is not up to date". I thought a hallmark of HA systems was fast failover? If I come to your house and knock on the door, but it takes you 10mins to get off the couch to open the door, it's perfectly acceptable for me to claim you were "unavailable". Pedants will argue the opposite.
how is this possibly a selling point?
> how is this possibly a selling point?
Context.
In an overview doc? The informal version is accessible to more audiences.
As the canonical design doc for the system? It's certainly not.
FWIW the phrase seems to come from here: https://bloomberg.github.io/blazingmq/docs/architecture/elec...
The full sentence reads:
"This section explains the leader election algorithm at a high level. It is by no means exhaustive and deliberately avoids any formal specification or proof."
Which makes a lot more sense.
And as noted in a sibling comment there is actually a TLA+ spec for the leader election: https://github.com/bloomberg/blazingmq/tree/main/etc/tlaplus
> Just like BlazingMQ’s other subsystems, its leader election implementation (and general replicated state machinery) is tested with unit and integration tests. In addition, we periodically run chaos testing on BlazingMQ using our Jepsen chaos testing suite, which we will be publishing soon as open source. We have also tested our implementation with a TLA+ specification for BlazingMQ’s elector state machine.
Edit: I did see the comparisons page, but.. well, there's more to life than Rabbit and Kafka!
Well, you know, there is the "Comparison With Alternative Systems" page[1]. :-)
Sadly it doesn't include NATS in the comparison. I would have been interested to see that. But the usual suspects (Rabbit & Kafka) are given a mention.
[1] https://bloomberg.github.io/blazingmq/docs/introduction/comp...
That said, there are indeed plenty of commercial and noncommercial alternatives, so now that it's been published I'm sure there will be some thorough people making comparisons soon!
RabbitMQ was probably avoided because nobody wanted to learn Erlang.
Also, R&D had a lot of experience building message oriented middleware "from scratch" in a low overhead high availability way, so first instinct was probably to start hacking in C++.
Nowadays it might be the case that some teams within Bloomberg need the performance or would rather have a bespoke solution instead of spending on migrating to something else off the shelf.
Keep in mind that this is a company that has its own implementation of most of the C++ standard library.
They have very strict standards which are necessary given that C++ really is a minefield. We did a great job avoiding UB by following those coding standards. There is some verbosity that can be avoided but hey these are details.
Personally, I have moved to Rust since then and I am not looking back.
I know Solana does mutual TLS over QUIC authentication between peers and does bandwidth rate limiting via stake. Also, the topology is self organizing. Is it worth copying this tech into a message queue system?
Maybe I'm old school but 99% of the use cases I've worked with in my 20 year career could scale with using a MySQL database table acting as a queue and some scripts querying the table and doing work. I have worked with maybe a dozen queuing technogies (including super expensive cloud Amazon SQS) and unless your app needs sub second latency on processing messages at greater than 5k per second I just don't see why using these sophisticated systems have benefits?
Hope this comment doesn't come across cynical or dismissive. I love seeing new tech like this come out and I am not saying that there aren't use cases I am sure there (maybe HFT?) but curious if anyone has case studies to share of legit queue systems that couldn't scale.
Having said that, you would need to match the protocol to the interaction style. Not all are appropriate.
I don't see anything that implies as such, all FAQ and comparisons to other message queues say nothing of security (I skimmed and ctrl+f so I could have missed it). This doesn't bode well for a network-based application written in a memory unsafe language.
edit: I'm surprised so many people are interpeting this as a weird cargo cult thing. Application security includes a lot of mundane and critical things like:
Authentication, authorization, message signing / authentication (e.g. HMAC), encryption, secret/key management, how the system handles updates, etc.
You can read a bit more about it here: https://en.m.wikipedia.org/wiki/Application_security
Your false dichotomy also implies that “unsafe” blocks are as unsafe as C++, which is not true. “Unsafe” in Rust turns off very few checks[0], most of them are still active. No one would write a serious Rust program entirely in unsafe anyways.
Regardless, asking about security considerations is a valid thing to do, even if it were written in Rust. Security is not just about memory safety.
[0]: https://doc.rust-lang.org/std/keyword.unsafe.html#unsafe-abi...
Having parsers (and serializers) proved absent of runtime errors (e.g. with something like SPARK) is a form of guarantee I wish I'd see as the default in any 'serious' network-facing library or component. It's not even that hard to achieve, compared to the learning curve of the borrow checker.
Once rust gets plugged into Why3 and gets some industrial-grade proof capacity, the question of 'is it written in rust?' will be automatic (as in 'why would you do it any other way?').
Also, it's a super bad-faith argument to talk about an entire program being in an unsafe block. Rust's unsafe is about specifically and intentionally opting out of certain guardrails in a limited scope of code where it should be fairly easy for a human reviewer to understand and validate the more complex invariants that aren't enforceable by the compiler itself.
"you can give up guaranteed safety in exchange for greater performance "
I mean, I don't often see documentation for networked services that goes in-depth on how exactly they write memory safe code in C or C++, as if that would mean the application is therefore secure; it's always much higher level.
I'm personally interested in both, but from a docs perspective I'd mostly expect the latter, more "security for users of this system" than anything else. I'm a little embarassed I didn't remember it until now, but I think the general term is "application security."
I'm concerned that I don't see any hints of either in the project. I'm not at a computer with access to my Github account so I can't easily search the source code for hints of obvious signs of care or neglect with respect to security.
Not sure where they are on the roadmap but authentication in particular is 'mid-term' but lacks detail.
I expect that if you want to deploy this then you're doing it on an internal network; its deployment examples are based on Docker so I expect it's relying on what you can do with k8s.
It looks like it's an after thought but at least on their mind now, which is very fair with respect to Bloomberg's wants/needs. It'd be nice if they had a bit of a warning about using this until it has some basic auth(n) and TLS since they're releasing it to the public. I think it is, relativley speaking, rude to release insecure networked software without giving users a notice as to what sorts of insecurity is at least known/expected.
And if you believe that, you're wrong
There's a hefty load of cargo cult mentality revolving the topic of security. Security is not a product, but a process. This sort of talk is getting tiresome.
This may be perfectly safe... however, I'd give it a few years of battle testing before touching it.
Here's an example:
https://github.com/bloomberg/blazingmq/blob/main/src/groups/...
Also they do make heavy use of managed pointers in the codebase.
Just because software is heavily used and deals with finance, doesn't mean it's secure or well written. Consider the Equifax breach of 2017, everyone's credit info leaked because Equifax was using a crusty old version of struts with known critical CVEs.
BlazingMQ also appears to be more of a "platform" or "service" that an app can use (sort of like Oracle, say) -- ZeroMQ includes libraries that one can use to build an app, service or platform, but none is provided "out of the box".
Which makes it harder to get started with ZeroMQ, since by definition every ZeroMQ app is essentially built "from scratch".
If you're interested in ZeroMQ, you may want to check out OZ (https://github.com/nyfix/OZ), which is a Rendezvous-like platform that uses the OpenMAMA API (https://github.com/finos/OpenMAMA) and ZeroMQ (https://github.com/zeromq/libzmq) transport to provide a full-featured network middleware implementation. OZ has been used in our shop since 2020 handling approx 50MM high-value messages per day on our global FIX network.
Finally! Thank you!
the github shows its codebase is in c++ with gcc-7.3, that means it might be using c++14 in the code, or even c++17, which is nice.
https://bloomberg.github.io/blazingmq/docs/architecture/netw...
Arguably, if you start with UDP but you also need to ensure retransmission and reordering, you eventually end up reimplementing TCP but poorly.
That guy really knows his stuff regarding performance.
Amazing to see a non-Rust product posted to HN.
> BlazingMQ is an actively developed project and has been battle-tested in production at Bloomberg for 8+ years.
> Carefully architected and written in C++ from the ground up with no dependency on any external framework, BlazingMQ provides…
What a weird selling point.
There are pros and cons to having no dependencies. Not long ago, it was a common decision for C++ projects because dependency management was a mess. But that hasn’t been the case for a long time (arguably we have the opposite problem - there are so many ways of doing it: Conan, vcpkg, bazel, spack, raw cmake, nix, etc)
So what would the pros and cons be today and why is it a selling point?
For instance, a pro is that everything is bespoke to the task at hand, no hammering square pegs into round holes designed for a slightly different use case.
A con is that the entire attack surface is managed by one team. CVEs identified and solved on another project is a pretty good thing when you depend on that project.
The main reason I’m surprised, though, is that there are some no-brainer dependencies these days. Fmt, catch2/gtest, metaprogramming libraries, etc.
This alone is an important point to avoid dependencies, specially if you're anything like Bloomberg.
Though my approach is to fork and maintain the fork until my PR is accepted, merged, and in a tagged release.
Another is to apply a diff to the dependency within the build system, which is what I do a lot in Nix (mainly to solve cmake issues on hermetic builds).
Either way, it isn't unsurmountable. It doesn't really matter so much if it's used as an executable rather than as a library, though in the case of the latter it's really handy to be able to e.g. pass a custom spdlog sink for logging.
I can definitely see why that's a bonus. So many data pipeline tools, but so many written in Java. It's for that reason that I've always reinvented the wheel. Perhaps I'll check this out...
I'm not saying they pulled it off because I would need to evaluate it, but having your own dependency stack can result in a much more streamlined solution for extra performance since you don't have to compromise to someone else's data model. As this is a high speed message bus, it's entirely appropriate.
These are also basically application-level network routers which really don't need huge levels of dependencies so to be honest I'd be running away from anything that claims it's written using 1000s of the newest, latest and greatest whatevers. Churning out a faster string grinder is what I'm looking for.
Seems situation dependent.
But what I think the other post is alluding to is the fact that in-order delivery is often not a requirement, and can be an anti-feature. I know that in every use of Kafka I have personally encountered the relative order of messages was quite irrelevant but the in-order implementation of Kafka converts the inability to process a single record into an emergency.
That is decent but still below the speed of the network, which should be the main blocker.
You likely need to read about message queues, why they are used, and why they have performance constraints. Message queues are a system, and the system to maintain those queues and content has an overhead. The project is building a solution that is higher performing than comparable systems.
Queuing things often makes things faster (!) as you are usually dealing with limited resources (ex: web and job workers) and need to operate on other systems (ex: payments) that have their own distinct capacity.
If you have 1000 web workers, and your upstream payment provider supports 10 payments per minute, you'll quickly have timeouts and other issues with 1000 web workers trying to process payments. Your site won't process anything else while waiting for the payment provider if all web workers are doing this task.
A queue will let you process 1000 web requests, and trickle 10 requests to your payment provider. At some point that queue will fill, but that's a separate issue.
Meanwhile, your 1000 web workers are free to process the next requests (ex: your homepage news feed).
Would you recommend something to read about that?
You'd usually want your queue system to fail the enqueue if it is full, and you'd want monitoring to ensure the queue isn't growing at an unsustainable rate.
It also forces you to think a bit about your message payload (rich data or foreign keys the worker loads).
RabbitMQ, Redis-based queues (ex: ActiveJob or sidekiq), Gearman, and others will all offer different mechanisms to tackle full queues.
- you need to ensure that your producers don't produce at a faster rate than your consumers can keep up
- if the consumer is unable to handle the request now, queueing it and handling it later is often not the right thing to do. Most likely by the time the resource is available the nature of the request will change.
TCP itself is already a queue. All these message queue systems also make the silly decision of layering their own queueing on top of TCP, leading to all sorts of problems.
Basically those things only sort of work if you have few messages and are happy with millisecond latency.
> Basically those things only sort of work if you have few messages and are happy with millisecond latency.
Not really... queues are great to defer processing or have longer running job than your HTTP and TCP timeouts will allow. Building large data export won't happen within a single HTTP request.
ex: https://shopify.engineering/high-availability-background-job...
1. They're referring to the queueing itself. They're not saying "this will make some other thing fast" they're saying "this is fast at queueing".
2. Queuing is not the opposite of making things fast.
TCP is just not a good fit for anything that needs high-performance or bounded sending times. What it's good for is giving you a simple high-level abstraction for streaming bytes when you'd rather not design your networked application around networking.
Building the right logic for your application is not difficult (most often, just let the requester resend, which you needs to do anyway in the age of fault-tolerant REST services, or just wait for more recent data to be pushed) and easier than tweaking TCP to have the right performance characteristics.
UDP also has very small message (datagram) size limits, so you would also need to implement some way to fragment and re-combine messages in your application.
At this point you've built an ad-hoc re-implementation of 80% of TCP.