HNHacker News
TopNewBestAskShowJobs

pmcf

140 karma · joined October 8, 2022

submissionscomments
pmcf··on Cassandra at Apple: 1000s of Clusters, 300k Nodes, 100 PB
That blog was posted 9 years ago. Needless to say, a lot of improvement and engineering has happened since then. The limited use cases of LWTs will soon be replaced with general, ACID transactions.

https://thenewstack.io/an-apache-cassandra-breakthrough-acid...

pmcf··on Cassandra at Apple: 1000s of Clusters, 300k Nodes, 100 PB
On the slide shown, it says "1000s of applications."
pmcf··on Cassandra at Apple: 1000s of Clusters, 300k Nodes, 100 PB
In my (extensive) experience in infrastructure. When people say the JVM is the problem, the JVM is never the problem. It's usually just a symptom of something else and lazy ops people just want to blame something and throw up their hands. I've never had to "tune a JVM" to make things work.

I'll give you an example in Spark. We had a huge job that was failing and after checking the logs, it was when results were being spooled to disk. More log diving showed a lot of GC on every node. At that point we could have gone down the route of tuning something in the JVM, but more digging found the real culprit. IOstats when the jobs ran showed while reading data, writes were completely blocked and write latency was in the 100s of ms. The spark executor trying to dump data was blocked and the first symptom that things were falling apart was... GC. We changed the scheduler on the nodes and magically everything worked great. The VM in JVM is virtual machine. Same rules for resources apply and if you run out of resources, don't expect the mythical ops faeries to save you.

pmcf··on Cassandra at Apple: 1000s of Clusters, 300k Nodes, 100 PB
I know this is HN and people love to troll, but it makes me sad when I see you using falsehoods to take a steaming dump on a group of engineers that are obsessed with building a database that can scale while maintaining the highest level of correctness possible. All in Open Source... for free! Spend some time on the Cassandra mailing list and you'll walk away feeling much differently. Instead of complaining, participate. Some examples of the project's obsession with quality and correctness:

https://cassandra.apache.org/_/blog/The-Path-to-Green-CI.htm...

https://cassandra.apache.org/_/blog/Finding-Bugs-in-Cassandr...

https://cassandra.apache.org/_/blog/Testing-Apache-Cassandra...

https://cassandra.apache.org/_/blog/Introducing-Apache-Cassa...

pmcf··on Cassandra at Apple: 1000s of Clusters, 300k Nodes, 100 PB
It’s a distributed system and if you have been a DBA for a single system like Oracle or MySQL there is a lot of new competencies to learn. That being said, completely doable and it’s typical to see small teams running massive amounts of Cassandra. At the same conference, Bloomberg talked about their large Cassandra footprint with only 4 people. If you want to run Cassandra in K8s there is the K8ssandra project that automates a lot. It’s a fast growing project as a result. (http://k8ssandra.io) If you want to use Cassandra and not run it, http://astra.datastax.com. One click and a few seconds, you get a completely serverless version of Cassandra that you only pay for what you use. I'm sure we will hear a lot more of these stories at Cassandra Summit in March (http://cassandrasummit.org)
← PreviousPage 2 of 2