YCSB Benchmark Results for YugaByte and Apache Cassandra
forum.yugabyte.com
forum.yugabyte.com
Also I highly doubt you want to have 30G heap with 60G RAM available. Having a larger heap most likely doesn't help you at all. It will just eat away memory you could use for mmapping or off heap things and increase GC pause length.
(This is the bane of running JVM any apps. You really need to know what you are doing with the GC. Also with apps like Cassandra there is no silver bullet. Just some general guidelines to follow and slowly adjust your values and monitoring of the changes as the best possible settings change with the load/hardware setup)
It would be nice if some of these were the defaults as opposed to some settings folks have to change - say by running a "./configure" or some such.
In this case, for example, it felt a little weird that the machine type which had enough cores had 60GB RAM, and giving only around 10-20GB to Cassandra "felt" wrong, but turns out the exact opposite is true! YugaByte is in C++ and can handle much larger heaps.
File system cache, in process decompressed page cache, key and row cache, off heap memtables etc. There are several different mechanisms for using memory that isn't on the garbage collected heap.
^ This is not a qualitative assessment of how Cassandra uses memory.
The JVM cannot know in advance if you prefer throughput or predictability.
https://www.google.pl/amp/s/www.infoworld.com/article/323539...
We have done the following p99 measurements: - a p99 comparison of YugaByte vs Cassandra using YCSB - ran ndbench for a week and looked at the p99 of YB
The results are looking really good for YugaByte in both cases. We are going to post that very soon to our forum, will x-post here as well.
In gist, we ran ndbench for 7 days and our p99 was under 6 ms with no "long tail". It is achieved by cautious architectural design and implementation choices. More details are available in our post.
In our new benchmark run, we set up Apache Cassandra using G1 GC and a smaller heap size as suggested. Still, we observed a multitude of differences in throughput and p99 latencies in YugaByte over Apache Cassandra.
> Rule 1: Developing a good DBMS requires 5-7 years and tens of millions of dollars.
> That’s if things go extremely well.
> Rule 2: You aren’t an exception to Rule 1.
> In particular:
> Concurrent workloads benchmarked in the lab are poor predictors of concurrent performance in real life. Mixed workload management is harder than you’re assuming it is. Those minor edge cases in which your Version 1 product works poorly aren’t minor after all.
Having seen the benchmarks for ScyllaDB vs Cassandra I think the crown easily goes to Scylla. That shouldn't be big surprise, they do have some amazing engineers on that team, including some of the creators of KVM. They wrote their own TCP/IP stack for Scylla and can run in kernel-bypass mode.
They also do some fascinating things like batching similar internal operations to get better instruction cache performance. I wouldn't have though I-cache would ever be a bottleneck for a database. It speaks to how highly-tuned Scylla is, where the other performance bottlenecks have been dealt with to such a degree that the I-cache behavior became a performance problem worthy of tackling.