New ScyllaDB Go Driver: Faster Than GoCQL and Its Rust Counterpart
scylladb.com
scylladb.com
> The big difference between our Rust and Go drivers comes from coalescing; however, even with this optimization disabled in the Go driver, it’s still a bit faster.
For anyone who's wondering, the Rust driver has coalescing support as of 9 days ago.
I just don't want to spin up EC2 instances manually, get the connections all working, make sure I can reset state, etc.
I already have a fork of Scylla where I removed a lot of unnecessary cloning of `String` but no way I'm gonna PR it without a benchmark.
I also opened a PR to replace the hash algorithm used in their PreparedStatement cache, which gets hit for every query, but they wanted benchmarks before accepting (completely fair) and I have none. `ahash` is extremely fast compared to Rust's default - https://github.com/tkaitchuck/ahash and with the `comptime` randomness (more than sufficient for the scylla use case) you can avoid a system call when creating the HashMap.
There are also some performance improvements I have in mind for the response parsing, among other things.
Not sure how decent a benchmark would be without running up servers in the cloud. So I guess provisioning infra would be a requirement?
So perhaps this could be run manually. But certainly possible
- Pulumi up infra - Run benchmarks - Collect results - Attach to PR.
1. Benchmarks of "pure" code like the response parser, which I could `cargo bench`. I may actually work on contributing this.
2. Some way to run benchmarks against a deployed server. I wouldn't recommend a Github action necessarily, a nightly job or manual job would probably be a better use of money/resources. If I could plug in some AWS creds and have it do the deployment and spit out a bunch of metrics for me that'd be wonderful.
I don’t have the exact numbers on me right now but I can share them tomorrow (along with the benchmark code) if you’re interested.
`ahash` has some good benchmarks here: https://github.com/tkaitchuck/aHash/blob/master/FAQ.md
The t1ha crate also hasn't been updated in over three years so the benchmark in this link should be current.
[1] https://github.com/tkaitchuck/aHash/blob/master/compare/read...
Edit: if you really think tha1 is faster I would open an issue on the aHash repo to update their benchmark.
Thanks!
you might want to particularly look into murmur, spooky, and metrohash. I'm not exactly sure of what the tradeoffs involved are, or what your need is, but that site should serve as a good starting point at least.
I've been thinking about this lately.
I wonder if we could standardize a benchmark format so that you could automatically do the steps of downloading the code, setting up a container (on your computer or in the cloud), running the benchmarks, producing an output file, and making a PR with the output.
So developers would go "here's my benchmark suite, but I've only tested it on my machine", and users would call "cargo bench --submit-results-in-pr" or whatever, and thus the benchmark would quickly get more samples.
(With graphs being auto-generated as more samples come in, based on some config files plus the bench samples)
1. I've re-opened by hashing PR and I'm going to suggest that they adopt ahash as the default hasher in the future.
2. I've re-written my "reduce allocations" work as a POC. Another dev has done similar work to reduce allocations, we took different approaches to the same area of code. I'm going to try to push the conversation forward until we have a PR'able plan.
3. I'm going to push for a change that will remove multiple large allocations (of PreparedStatement) out of the query path.
4. Another two devs have started work on the response deserialization optimizations, which is awesome and means I don't have to even think about it.
I think we'll see really significant performance gains if all of these changes come in.
Disclaimer: I work for ScyllaDB, although not on drivers. I can forward your question to relevant people.
https://github.com/scylladb/scylla-stress-orchestrator/wiki/...
They consistently demonstrate that we are under using our CPUs compared to potential.
I'm not sure how much "low hanging" fruit there is, though. A lot of modern slowdown is architectural. Nobody sat down and thought about the flow through the system holistically at the design phase, and the system rolled out with a design that intrinsically depends on a synchronous network transaction every time the user types a key, or the code passes back and forth between three layers of architecture internally getting wrapped and unwrapped in intermediate objects a billion times per second (... loops are terrible magnifiers of architecture failures, a minor "oops" becomes a performance disaster when done a billion times...) when a better design could have just done the thing in one shot, etc. I think a lot of the things we have fundamental performance issues with are actually so hard to fix they all but require new programs to be written in a lot of cases.
Then again, there is also visibly a lot of code in the world that has simply never been run through a profiler, not even for fun (and it is so much fun to profile a code base that has never been profiled before, I highly recommend it, no sarcasm), and it's hard to get a statistically-significant sense of how much of the performance issues we face are low-hanging fruit and how much are bad architecture.
Is it along the lines of “we want to collect all the data we can in case we want to use or analyze it at some point”, or are there real use cases?
The key.
And I was wondering how can a tracing GC outperform a non-tracing-GC memory manager.
(One fast way to manage allocations is to use an arena allocator that allocates memory by incrementing a pointer, and frees memory all at once. This is pretty effective for simple, short-lived requests.)
In C++ you also need to minimize allocations, but it’s radically easier to do in C++ than in C#.
The cliché is that malloc/free style memory management has to touch all the garbage in order to free it, while a semispace GC only has to copy the live data once in a while. The garbage is ignored.
A GC can also let you use more efficient concurrent data structures - many sophisticated concurrent objects require a tracing GC for implementing correctly - which can improve the performance of your application code.
What sort of "sophisticated concurrent objects" are you thinking of?
https://medium.com/@tylerneely/fear-and-loathing-in-lock-fre...
To me "it could be more difficult without" and "requires" are quite different claims, especially in the context of what's possible and why.
I am not sure of the performance or implementation difficulty but the data structure seems to be what you are talking about.
An example would be using (TVar (HashMap k v)) in Haskell.
Hadn't heard of the pre-coalesce millisecond pile up technique.
Favorited, thank you, sincerely!
This is basically Nagling and/or TCP_CORK right?
There are some specific features like shard-aware queries and shard-aware ports that naturally won't apply. But they will work.
Now I want to know more. :-)
Your general comment is correct. I see it often with GPU algorithms which, no surprise, are also much faster on CPUs (using something like ISPC to compile them).
An approach that tends to be ignored by those rewrite X in Y.
ScyllaDB is a wide column store which is, in fact, a row store; you can call it a "key-key-value," since it had a partitioning key and a clustering [or "sort" key]. Which is more for transactional workloads [OLTP]. So it is more comparable with Cassandra or DynamoDB.
So they are really designed for different sorts of things.
That being said, ScyllaDB has some features, like Workload Prioritization, so you can run analytics, like range or full table scans against it without hammering your incoming transactions. But it wasn't designed specifically for that.
https://groups.google.com/a/chromium.org/g/chromium-dev/c/EU...