Going Head-To-Head: Scylla vs. Amazon DynamoDB
scylladb.com
scylladb.com
> Please keep reading to see how diligent we were in creating a fair test case
I hope this is true, but my cursory reading found several unfair spots that seem at odds with this statement.
> 3-node cluster on single DC | RF=3
DynamoDB works across AZs and will continue to be available even when a data center goes down. That cross-AZ operation has benefits as well as latency costs that are not present in the ScyllaDB setup.
> We hit errors on ~50% of the YCSB threads causing them to die when using ≥50% of write provisioned capacity
That is surprising. It is quite common to use your full provisioned capacity without problems. My guess is that there is something not ideal about the YCSB DynamoDB library or its configuration. I'm not familiar with YCSB: does it give you stack traces that indicate why the threads failed?
> Sadly for DynamoDB, each item weighted 1.1kb – YCSB default schema, thus each write originated in two accesses
This is what made me come write this comment. You specifically knew this was a pessimistic case which could be easily addressed to allow DynamoDB to operate at a lower cost. Is arbitrarily settling for the default 1.1kB item size on an artificial benchmark fair? Good engineering teams use their tools the way that gives them the most benefit. Calling out that you may have to work to ensure your use case doesn't have pessimistic characteristics would clearly be fair, but I'm not convinced that just picking arbitrary benchmark settings is.
1. Scylla works within different AZs too We are topology aware and can have as good and even better HA than DynamoDB. Our design is based on C*
2. We were surprised with getting small utilization too. As I wrote in the article, I think Dynamo had a hard time reaching to 1TB that quick. The population started fine and deeper in the the population it failed. Only a decrease in the rate solved it.
3. 1.1kb The is the default table setup by YCSB. It isn't optimal for Dynamo but that's life.. think about if you store a blob of 1kb - together with the key it will be more than 1kb. We were upfront about it and you can make your own calculation. Scylla will still be better.
The only use case where Dynamo is better in price is when you store lots of data but with a tiny IOPS reservation. We'll address this case over time too. Cheers!
Did your test use this feature of Scylla? As I understand the documentation, it doesn't appear so (maybe that needs to be fixed). If that's the case, do you think this is an example of being diligent in your fairness?
> Only a decrease in the rate solved it.
I don't think that is the only way to solve it. Something is probably wrong here, and I would expect that diligence to fairness would dictate trying to understand and correct that. This benchmark would be more compelling if problems like this weren't just ignored, but actually solved. Any team that cares about the cost of their database would do that.
> It isn't optimal for Dynamo but that's life.
Yes, life is unfair and this can happen. This test specifically claimed to be fair, though. Again, for this benchmark to be meaningful and compelling, this kind of thing should be addressed.
> The only use case where Dynamo is better in price is when you store lots of data but with a tiny IOPS reservation
This is a surprising statement. If this is the case, I highly recommend addressing the issues so that you are using DynamoDB like a typical engineering team would. A blanket statement like this with questionable benchmarks as evidence is just not compelling. It would be much more helpful to see where DynamoDB really can't stand a chance competing against ScyllaDB even when it's fully optimiized.
Are you ignoring maintenance costs? Do nodes get added and removed without paying an engineer? How do you keep the nodes up-to-date as far as security updates or adding new features? How much overhead cost is engineering on-call to respond when the cluster has issues (e.g. nearing capacity limits, nodes failing, loss of AZ availability)?
I'm also curious if you are only considering use cases that you typically already see on ScyllaDB and Cassandra. How about the use case where I have a load that increases without warning by 10x or where I need regular backups or that requires 5000 nodes?
While there are indeed deployments that large, I've yet to see a 5,000 node _benchmark_. Though if you've seen tests that big, please send me the link!
https://read.acloud.guru/why-amazon-dynamodb-isnt-for-everyo...
You might also want to read about the scaling behavior of DynamoDB On-Demand (https://docs.aws.amazon.com/amazondynamodb/latest/developerg...). Of course you can't immediately scale to any arbitrary request rate, but the way it's implemented sounds fine to me.
So do you have a source for your claim?
Since I see you're with AWS, you're welcome to help us debug this error 500 on population. I guarantee we'll update the blog if we'll see it works fine
I didn't get your surprise, I actually am 100% open about digging deep and come up with a uncommon use case that Dynamo can finally be a good alternative.
Believe me, I can come with examples that will put Scylla in a much better light.
We do offer a fully managed service on AWS with 4x-6x cost saving vs Dynamo where all the backup and auto scale is done by the Scylla team.
Are you suggesting that the metrics would not change? I assure you that is not the case: inter-AZ latency is 1-2 orders of magnitude higher than intra-AZ latency.
You'll need to re-run your benchmarks in a multi-AZ cluster in order to provide an apples-to-apples comparison.
The intra AZ latency according to AWS is sub millisecond! It's inline with what our customers see.
Since all of the numbers in the blog are above a millisecond, the result does not change.
Why don't you try it yourself and see the difference? Relying on some third party website is no substitute for testing.
As you can see, the ping latency is 0.5ms to 1ms on rare occasions. Of course that ping has a cost even within a single AZ so the above numbers reflect much more than the AZ round trip time.
You're welcome to add 1ms to our latency numbers, it's 3x-4x better in the 99th case so there's enough room...
64 bytes from 35.173.249.24: icmp_seq=14 ttl=63 time=0.584 ms 64 bytes from 35.173.249.24: icmp_seq=15 ttl=63 time=0.462 ms 64 bytes from 35.173.249.24: icmp_seq=16 ttl=63 time=0.848 ms 64 bytes from 35.173.249.24: icmp_seq=17 ttl=63 time=0.465 ms 64 bytes from 35.173.249.24: icmp_seq=18 ttl=63 time=0.451 ms 64 bytes from 35.173.249.24: icmp_seq=19 ttl=63 time=1.00 ms 64 bytes from 35.173.249.24: icmp_seq=20 ttl=63 time=0.854 ms 64 bytes from 35.173.249.24: icmp_seq=21 ttl=63 time=0.446 ms
Continues here: https://issues.apache.org/jira/browse/CASSANDRA-14448
Create the table with 4x the provisioned capacity and then dial down to your nominal levels after it's been created.
So you solve 500 errors by provisioning 4x your usage? Seems like solving a leak in the gas tank by instructing the user to simply fill up the tank more often.
Nice business if you are the gas station!
They really are. In order to make effective (let alone efficient) use of DynamoDB you need to keep in mind exactly how it works behind-the-scenes and exactly how the charging models works at all times. And you have very few, very coarse levers to pull (over-provision, under-provision, etc.)
Every in-depth explanation or tutorial of DynamoDB I've seen advises to do things like this in multiple different load scenarios. There's a pretty good intro for DynamoDB over at acloud.guru: https://acloud.guru/learn/aws-dynamodb - beware though, it costs money and it's very long (though it's split up so you can cherry pick the bits you care about).
I think DynamoDB might be worth investigating when you're hitting a scale where most other off-the-shelf options don't work and you'd be heading down the path of building your own solution on top of something anyway.
Seen your comment about LWT, agree.. We knownly chose the first make MV/SI great before committing to another big feature. Glad to say that with Scylla 3.0 release this month we'll finally get LWT done
I know this; it's why I was disappointed in the way this test was designed. It would be great to see as close as possible comparison between the two then mention improvements and trade-offs that you can still make (such as single-AZ or self-service) since you have different knobs to turn as well as DynamoDB quirks (such as item size boundaries) that can make ScyllaDB more attractive for your use-case.
Also, it sounds like Scilla has a managed solution too, that costs a fraction of dynamodb, at scale.
Dynamodb to me is the new generation IMS database from mainframe cobol times - fully proprietary, locked down, runs only on proprietary hardware, serviced by one vendor. It’s only a matter of time before companies using it will be stuck with multimillion $$ yearly bills, if they aren’t already.
Do you? At least with DynamoDB On-Demand there shouldn't be much maintenance left to do.
QUICK STRAW POLL: What is the right way for that information to be gathered and presented?
a. There is no bias. Companies can and should present research that may favor them, as long as they cite good sources.
b. An industry analyst, paid to research multiple companies and options and present that information to their respective (most likely paying) customers.
c. A journalist, presenting research in public, funded by advertising for unrelated interests.
d. Reporting by peers/actual users of the system on how the products compare, and possibly about how it aligns with their technical and business goals.
e. Trust no one. Conduct your own research, and publish it if you see fit to so do, or keep it to yourself.
I see potential "disinformation vulnerability" with each approach. How ought we get the best info on how to align our organizations with technologies?
- Here's a webinar by one of our customers explaining how they migrated from Mongo+Hive, received better consistency and simplicity while saving 5X: https://www.youtube.com/watch?v=1hXKd_rNyuE
- See how kiwi.com got a CRAZY gain with Scylla vs Cassandra: https://youtu.be/Bqh09LG_QDE?t=833
- See Grab (South east Asia Uber) about Dynamo vs Scylladb: https://www.scylladb.com/users/case-study-grab-hails-scylla-...
Go ahead and give it a try and see for yourself
Seeing a multi-AZ configuration compared to a single-AZ config is an immediate disqualifier, without getting into some of the other comments on item size disparity and YCSB quirks.
This was a really poor response from Scylla on these fronts: https://news.ycombinator.com/item?id=18677511
As you can see, the ping latency is 0.5ms to 1ms on rare occasions. Of course that ping has a cost even within a single AZ so the above numbers reflect much more than the AZ round trip time.
You're welcome to add 1ms to our latency numbers, it's 3x-4x better in the 99th case so there's enough room...
64 bytes from 35.173.249.24: icmp_seq=13 ttl=63 time=0.792 ms 64 bytes from 35.173.249.24: icmp_seq=14 ttl=63 time=0.584 ms 64 bytes from 35.173.249.24: icmp_seq=15 ttl=63 time=0.462 ms 64 bytes from 35.173.249.24: icmp_seq=16 ttl=63 time=0.848 ms 64 bytes from 35.173.249.24: icmp_seq=17 ttl=63 time=0.465 ms 64 bytes from 35.173.249.24: icmp_seq=18 ttl=63 time=0.451 ms 64 bytes from 35.173.249.24: icmp_seq=19 ttl=63 time=1.00 ms 64 bytes from 35.173.249.24: icmp_seq=20 ttl=63 time=0.854 ms 64 bytes from 35.173.249.24: icmp_seq=21 ttl=63 time=0.446 ms 64 bytes from 35.173.249.24: icmp_seq=22 ttl=63 time=0.460 ms 64 bytes from 35.173.249.24: icmp_seq=23 ttl=63 time=0.509 ms
1. running a ping test is a fine start but wouldn’t a more realistic test be to rerun the benchmark? 2. in a multi-az setup you’re going to have to pay for inter-az transfer costs between ec2 nodes. data transfer between ec2 and dynamodb is free.
Also on the pricing front reserved capacity really drops your costs with dynamodb if you know you’re going to be with it for some term of time. Not that that needed to be included in the benchmark but... aws pricing is complicated, to say the least.
I think this should be 10GB partitions.
By the way, maybe it’s worth mentioning adaptive capacity? https://aws.amazon.com/blogs/database/how-amazon-dynamodb-ad...