841 karma · joined April 18, 2019
I’m also not sure if DNA-seq data refers to the human host, or just all DNA they were able to sequence (which would include bacteria as well I guess?)
RE:numbers: https://www.warpstream.com/blog/warpstream-benchmarks-and-tc...
RE: comparing to a single-zone Kafka cluster. A lot of people really dislike operating Kafka. Some people don't mind it and that's cool too, but its not the majority in my experience.
You don't have to keep the data stored in S3 express one zone forever, you can just land it there and then immediately compact it to S3 standard. You still pay the higher fee to write to S3EOZ, but not the higher storage fee.
WarpStream does this, data gets compacted out within seconds usually. Of course this is now... tiered storage. But implemented over two "infinitely scalable" remote storage systems so it gets rid of all the operational and scaling problems you have with a typical tiered storage Kafka setup that uses local volumes as the landing zone.
I could make the hardware free for Kafka in the comparison, and WarpStream would still come out significantly more cost effective. Cloud networking is really expensive.
This is more or less exactly what WarpStream is: https://www.warpstream.com/blog/minimizing-s3-api-costs-with...
Kafka API, S3 costs and ease of use
https://docs.WarpStream.com is the best document we have right now. This sounds interesting though, can you jump in our slack or shoot me an email at founders@warpstreamlabs.com ? Happy to support you any way we can!
Yeah we’ve run into a number of people who’ve rolled their own solution in this space. The “push pointers to S3 through traditional Kafka” approach is a very practical one.
Was this memq at Pinterest, or something else?
That said we’d be stoked to get this working as an additional tool for people. Do you want to shoot me an email or join our slack so we can discuss further? I can probably prioritize whatever protocol features were missing to get it working.
founders@warpstreamlabs.com
FWIW we’re considering a version where you can host the metadata yourself for enterprise users. For the free tier though we didn’t think it made sense since for a workload that could fit into our free tier, it didn’t seem like anyone would want to be responsible for the metadata layer themselves. Would love your feedback on that.