Building a Distributed Log from Scratch: Sketching a New System
bravenewgeek.com
bravenewgeek.com
> we will focus on what it takes to build something like this using Kafka and NATS Streaming as case studies of sorts—Kafka because of its ubiquity, NATS Streaming because it’s something with which I have personal experience.
Why not just set everything into essentially an S3 bucket, and then process off of it. What's the benefit of a broker in between them really?
Dumping everything into an S3 bucket may be perfectly fine and even efficient for some projects. I personally don't see myself wanting to use S3 buckets for transmitting messages between multiple different systems with different purposes.
If you have a lot of different services producing data that multiple services will consume, then a pub/sub system can simplify things. Some systems, like Kafka, also have added redundancy/scaling benefits.
Currently this is done for the control plane by just loading them directly into Elasticsearch. Increasing the availability for all users, the consideration is, does a pub/sub add anything that just throwing the data directly into an s3 like bucket (swift in our case).
I just assumed it did. Everyone does it. Scale, decoupling the sender from the receiver...
But... I mean, why? What's the point of an additional hop, and another thing to manage. If your source is already wrapping metadata, and perhaps putting it into a structure, what is the value of a pub/sub there? If you can load it to the pub/sub, you can load it to the s3 bucket in the same manner.