I was a bit surprised to see the announcement about making Redpanda open source (although very interesting for poking around!).
Question: are you, as a company (Vectorized) pursuing the same business model as the Confluent one? OSS Kafka, but paid official support, cloud, connectors and registry?
I think there is a big shift when you have a single binary. i.e.: no one really complains from running ngix, etc because it's easy to get up and running.
so the gist we wanted to let everyone use it and reserve the right to be the only hosted redpanda provider.
i.e.: think GDPR compliance as long as your data streams are in JSON or some predefined format.
WASM basically allows us to push computational guarantees to the storage engine itself without a separate cluster.
I've failed to grasp if it's an "alternative plugin-engine" kind of system to extend Redpanda, or you're storing data as WASM and therefore it's executable (to take your example: GDPR compliant auto-expire if it's past a certain date).
It seems totally fair to me. Good luck with this approach. I think yours odds are good.
This is simply not true. The restriction applies to 100% of everyone using the software under the license.
If an alternative service provider can't legally host the service for me, I am restricted from selecting an alternative vendor if my needs converge from the available vendors offerings.
Further, it still isn't open source.
1. You can't just take this and make money from it
2. They don't accept or expect external inputs.
Its the same as cockroachdb in that regard
If you have any questions or think your use may be confusing please reach out to us.
Looks like the project is mostly written in C++ and Go. What was the reason for this choice? Have you considered other languages, like Rust, Zig or similar instead of C++? TBH not sure what an alternative to Go would be, maybe JVM AOT-compiled with NativeImage (but AFAIK that's still experimental).
Did Go's GC and/or C++'s lack of GC help/impede the project? IMO memory management is one of the main differences between languages... the other is concurrency / memory model / undefined behavior, where JVM is significantly ahead of the rest (no undefined behavior), I'm not sure exactly where Go stands (there seems to be a memory model, but no mention of undefined behavior or lack thereof).
Sure thing. I basically was playing with an RPC framework to make my old company project fast in 2017 see github.com/smfrpc/smf
I wanted to play with dpdk then.
When I started this in 2019 I wanted to use a framework that was battle tested, so seastar.io fit the bill.
There is also a big part of it that I had been professionally programming in c++ for many many years.
Last, this evolved from a prototype on my laptop in miami to company.
I think rust would he am excellent choice as well and I bet we'll write a bunch of rust in the not too distant future.
As a satisfied Scylla convert, I'm looking forward to trying Redpanda.
Main departure is we are only API compatible. It was an explicit choice to use Raft vs ISR and to not use ZK, etc.
but indeed, seastar is a really fun framework to build storage systems in. I know ceph is also doing a re-write of a subsystem in seastar for example.
there is no substitute to testing tho
there are 2 levels here. 1) raft has a proof (and a great phd dissertation from diego), but what matters is if we actually implemented it correctly. so 2) is we need to continuously test it. Denis did a lot of similar work at CosmosDB (microsoft) and has spent his career working on consensus.
Hopefully these eases some concerns.
I'm asking partly because we must be able to offer closed on-prem installations as well as SaaS on a cloud. I'm looking for a low-ops component that will not fail me (as often as alternatives would) :)
I'd be interested in the write amplification since you went pretty low level in your IO layer. How do you guarantee atomic writes when virtually no disk provides guarantees other than on a page level which could result in destroying already written data if a write to the same page fails - at least in theory - and so one has to resort to writing data multiple times.
Seems like a lot of questions on here are just: how much better is the performance and why?, so maybe showing is better than telling.
Best of luck!
Though what is to me more interesting, is what happens when you inject failures while the benchmarks are running.
I'm a big fan of kafka as an abstract building block, but not so much the actual implementation, which is as painful to setup as a consultancy-based business model might make you suspect it would be, especially if you need reliability. The other problem is that performance kind of sucks, apart from potential latency spikes due to GC pauses I found even the average latencies for reliable end to end (in a fast local network and on decent sized hardware) not in right order of magnitude ballpark.
-Redpanda (free) - comparable features to Kafka, no limits -Redpanda Enterprise (paid) - Additional features (security, WASM, tiered storage, support etc) -Vectorized cloud (Free and paid tiers) - Hosted in AWS+ GCP
In general it probably best to run with:
--smp <n> --overprovisioned
for a container with <n> CPUs.
These are standard Seastar flags.
we also have a container image `vectorized/redpanda:latest`
I think people LOVE the kafka _api_ but they have a hard time operating clusters at scale. So we decided to keep the same API but solve the problem of operational complexity.
That is very true, and this is what you should emphasise. Wish you the best of luck!
The actual speed of Kafka has rarely been a concern (but huge numbers of partitions are, which makes rebalancing a pain) in my experience, in fact it was mostly overkill. But operational complexity was definitely an issue!
Without scrolling, all the text thats displayed is this:
"Redpanda
A Kafka® API compatible streaming platform for mission-critical workloads.
Try Redpanda today"
It tells me what it is, which is good. It does not tell me why it is better, for that I have to scroll. Kafka is already suited for mission-critical workloads, so that is not a unique value proposition. Maybe:
"Redpanda - 100% Kafka API compatible, but without the headaches. Forget Zookeeper, forget rebalancing issues. Instead, enjoy reliable message delivery, 10x faster speed and ultra-low latencies due to our thread-per-core architecture."
Something like that, plus a visible "call to action" button, maybe "try it out" or "download".
Could also think about a pretty graph comparing latencies or smth. People love pretty graphs.