HNHacker News
TopNewBestAskShowJobs

richieartoul

841 karma · joined April 18, 2019

richardartoul gmail
submissionscomments
richieartoul··on S3 Is the Future, S3 Is the Past
Nice article. I agree that it is a bit of a shame that everything is forced to be so S3-centric (and I say that as someone whose work helped motivate a lot of people to do that), but right now its an unfortunate reality of running software in the cloud because cloud networking and SSDs are so expensive that you really are required to use S3 if you want a system that can handle "big data" scale workloads cost effectively
richieartoul··on PlanetScale for Postgres is now GA
It's not a landing page, it's a blog, and if you read the first few sentences of the post it becomes immediately clear what service PlanetScale provides.
richieartoul··on 'Obelisks': New class of life has been found in human digestive system
The paper says: “in companion DNA-seq data from this project, no detectable Obelisk reads are found” which I think is getting at what you’re asking, but I’m not sure if my understanding is correct.

I’m also not sure if DNA-seq data refers to the human host, or just all DNA they were able to sequence (which would include bacteria as well I guess?)

richieartoul··on Tiered storage won't fix Kafka
Short answer is that MSK has (almost) all the same problems as OSS Kafka. Kinesis streams is a different beast that would require its own blog post.

RE:numbers: https://www.warpstream.com/blog/warpstream-benchmarks-and-tc...

richieartoul··on Tiered storage won't fix Kafka
I talk about this more here: https://www.warpstream.com/blog/s3-express-is-all-you-need

RE: comparing to a single-zone Kafka cluster. A lot of people really dislike operating Kafka. Some people don't mind it and that's cool too, but its not the majority in my experience.

richieartoul··on Tiered storage won't fix Kafka
(WarpStream co-founder here)

You don't have to keep the data stored in S3 express one zone forever, you can just land it there and then immediately compact it to S3 standard. You still pay the higher fee to write to S3EOZ, but not the higher storage fee.

WarpStream does this, data gets compacted out within seconds usually. Of course this is now... tiered storage. But implemented over two "infinitely scalable" remote storage systems so it gets rid of all the operational and scaling problems you have with a typical tiered storage Kafka setup that uses local volumes as the landing zone.

richieartoul··on Google made me ruin a perfectly good website (2023)
Is there a word for things that are both hilarious and tragic? I laughed out loud multiple times. Kudos to the author for making such a depressing topic so hysterical.
richieartoul··on Distributed coroutines with a native Python extension and Dispatch
Storing the data in the customer's S3 bucket is a really good idea IMO. Means they just have to trust you to run the coordination reliably, but not necessarily to keep their data secure which is a much easier bar to clear.
richieartoul··on Deterministic simulation testing for our entire SaaS
I think I just followed the official recommendations I found (which are probably stale now). I'll update it to r5, but it doesn't really matter. The price difference between the two is like 5%, but hardware only ends up representing a tiny fraction of Kafka's cost at scale (the real cost comes from EBS and inter-zone networking).

I could make the hardware free for Kafka in the comparison, and WarpStream would still come out significantly more cost effective. Cloud networking is really expensive.

richieartoul··on Deterministic simulation testing for our entire SaaS
https://www.warpstream.com/blog/rss.xml
richieartoul··on Differential storage: A key building block for a DuckDB-based data warehouse
Makes sense, very cool. Thanks!
richieartoul··on Differential storage: A key building block for a DuckDB-based data warehouse
This is pretty cool. I'm a bit of a noob when it comes to stuff like FUSE. There is a bunch of commentary in the blog post about how when DuckDB does X, the Differential Storage implementation does Y. If the differential storage implementation just exposes itself as a filesystem though, how does it know what DuckDB is doing at the application layer? For example, how does it know from the filesystem layer that DuckDB is performing a snapshot?
richieartoul··on S3 Express Is All You Need
Do you have any more details you can share about the performance of EFS? I've never met anyone who has actually used it in anger.
richieartoul··on Show HN: Gogosseract, a Go Lib for CGo-Free Tesseract OCR via Wazero
This is awesome and one of the things I’m really excited about with WASM, and specifically Wazero. The Wazero team is top notch. Now someone just needs to do this with zstd and make it fast…
richieartoul··on OpenTelemetry at Scale: Using Kafka to handle bursty traffic
(WarpStream founder)

This is more or less exactly what WarpStream is: https://www.warpstream.com/blog/minimizing-s3-api-costs-with...

Kafka API, S3 costs and ease of use

richieartoul··on OpenTelemetry at Scale: Using Kafka to handle bursty traffic
You have to do a bit more than that if you want exactly once end-to-end (I.E if Kafka itself can contain duplicates). One of my former colleagues did a good write up on how Husky does it: https://www.datadoghq.com/blog/engineering/husky-deep-dive/
richieartoul··on How FoundationDB works and why it works (2021)
It runs fine in AWS, Snowflake and many others run it there. The most recent FoundationDB paper goes into a lot more detail on their recovery protocol, it’s a lot more nuanced than you think, but it works extremely well
richieartoul··on How FoundationDB works and why it works (2021)
Newer versions are moving towards a custom btree storage engine called Redwood
richieartoul··on How FoundationDB works and why it works (2021)
Replied above
richieartoul··on How FoundationDB works and why it works (2021)
Requests to the sequencer are batched heavily. If the sequencer fails, the cluster goes through a recovery and will be unavailable for 2-3 seconds and then recover.
richieartoul··on Jacobin: A more than minimal JVM written in Go
I know the link stresses that correctness and code clarity are the primary goals of this project, but I’m curious if performance is a goal as well? Do you hope that people will run “real” workloads with this or use it to embed other software in their Go applications?
richieartoul··on Jacobin: A more than minimal JVM written in Go
This is so cool. Can’t wait to go digging around through the code this weekend. Good luck with the project!
richieartoul··on The broad set of computer science problems faced at cloud database companies
I’ve built pretty much my entire career around this problem and it still feels evergreen. If you want a meaningful and rich career as a software engineer, distributed storage is a great area to focus on.
richieartoul··on The broad set of computer science problems faced at cloud database companies
You usually have to redesign the storage from the ground up around S3. Some databases do transparent data tiering with S3 which can work ok too depending on the use case.
richieartoul··on Kafka is dead, long live Kafka
(WarpStream Cofounder)

https://docs.WarpStream.com is the best document we have right now. This sounds interesting though, can you jump in our slack or shoot me an email at founders@warpstreamlabs.com ? Happy to support you any way we can!

richieartoul··on Kafka is dead, long live Kafka
It’s getting late and I’m not 100% confident I’m sure what you’re asking, but I believe the answer to your question is yes. When you produce a message/batch you get back offsets for each message you produced. If you produce again after receiving that acknowledgement, the next set of offsets you receive are guaranteed to be larger than any of the offsets you received from your previous request.
richieartoul··on Kafka is dead, long live Kafka
(WarpStream co-founder)

Yeah we’ve run into a number of people who’ve rolled their own solution in this space. The “push pointers to S3 through traditional Kafka” approach is a very practical one.

Was this memq at Pinterest, or something else?

richieartoul··on Kafka is dead, long live Kafka
Our kafka protocol support is still incomplete. We’ve documented our progress here: https://docs.warpstream.com/warpstream/reference/kafka-proto...

That said we’d be stoked to get this working as an additional tool for people. Do you want to shoot me an email or join our slack so we can discuss further? I can probably prioritize whatever protocol features were missing to get it working.

founders@warpstreamlabs.com

richieartoul··on Kafka is dead, long live Kafka
DMd you on twitter
richieartoul··on Kafka is dead, long live Kafka
(WarpStream Cofounder)

FWIW we’re considering a version where you can host the metadata yourself for enterprise users. For the free tier though we didn’t think it made sense since for a workload that could fit into our free tier, it didn’t seem like anyone would want to be responsible for the metadata layer themselves. Would love your feedback on that.

Page 1 of 4Next →