This sounds like a kafka-like streaming system, but backed onto s3-objects?
Doesn't this mean that write performance is going to be bad?
This sounds like a kafka-like streaming system, but backed onto s3-objects?
Doesn't this mean that write performance is going to be bad?
And it's also fair to question write performance, since it's backed by object storage. The optimization is primarily from the shared WAL across streams, server-side batching, and client-side in-memory pipelining, especially with HTTP/2, without as much connection pool overhead.
In practice, you can go to the extent of achieving up to 100 MiB/s throughput per stream. Considering how granular streams can be, you'd rarely need as much. The latency for a durability ACK is, however, the price to pay, which is going to be ~250 ms, or lower with S3 Express, which I'd say covers most real-time use-cases. The design itself is easy enough to extend to a disk-staged WAL for single-digit durability ACK latency.
I'll be setting up a GCP deployment example similar to AWS soon. I'll be sure to try Rapid Bucket as well, thanks for sharing!
I would also check out Tigris [1], which has an S3-compatible API. Their main claim to fame is that buckets are low-latency, multi-region and replicated by default, so supposedly you get region-local latency no matter where you are reading or writing from. I have not done any rigorous performance comparisons, though. What's amazing, if it does perform well, is that egress is free, and the pricing is otherwise the same as GCS/S3.
There are plenty of usecases for lower scale or higher latency (my examples are somewhat unique), and owning the opinionated middle instead of claiming to cover everything is a really useful thing, but acknowledging that the system is opinionated such that it covers a specific set of things well is generally a better argument than 'this basically does everything that people need'.
Where Kafka starts to fall short is routing. If you want to access the data of one user from user-events-topic, that’s expensive to do. Most other streaming technologies are built around the same design, such as Kinesis.
There are other implementations that support the Kafka wire protocol and are cheaper in exchange for latency, e.g., AutoMQ and WarpStream.
That said, I’ll release Disk/EBS-staged WAL soon enough: https://github.com/PicoMQ/picomq/issues/13 as an add-on to cover low-latency needs.
I had heard people talk about the operation pain involved in keeping Kafka alive (which is a thing for sure), but what I was surprised by was how many things behaved in a slightly unobvious manner that wasn't loudly-documented (e.g. if you're using transactions for RWP loops the default rebalance protocol is unsound and transaction markers take up an index in the log so you no longer have contiguous indices in your message stream etc).
But because the Pico's semantics are close to Kafka, but not tightly coupled, it's quite feasible to support the Kafka protocol, * with some restrictions *, such as a topic can only ever have one partition, no support for transactions (since they would only really apply when producers are publishing to multiple streams), and more along those lines. As a result, you'd get Kafka with no topic tax and a lot less operational complexity.
Regular S3 has write latency 100-150ms, which might be fine depending on your workload anyway.