If you are going to benchmark a distributed system, you really need to set up more than 1 server.
(Disclaimer - work at Datastax)
If you are going to benchmark a distributed system, you really need to set up more than 1 server.
(Disclaimer - work at Datastax)
I think what they meant with "1000 nodes" is that the dataset they're using for the benchmark is synthetic monitoring data (where the thing being monitored are servers).
And the way they generated the synthetic data set is by having 1000 imaginative servers produce one sample per second, (i.e. have a script that writes out 1000 * duration_in_sec fake samples -- I believe this is the code that does it https://github.com/influxdata/influxdb-comparisons/tree/mast...)
Posting 1 node benchmarks of distributed databases seems suboptimal.
I am under the impression that Cassandra's performance comes from its distribution capabilities.
In fact for most dev I use 3 nodes on my laptop, and most of our "unit" tests are multi-node as well (closer to integration tests by most measures).