A Bunch of Money on AWS and Some Benchmark Results
memsql.com
memsql.com
1. Total run time is not an appropriate way to summarize the performance across queries, because some queries take 100x longer than others. The appropriate way to summarize this kind of data is to use the geomean [3].
2. The official TPC-DS queries make heavy use of grouping sets, which are a rarely-used SQL feature. I think TPC-DS is better if you rewrite the queries to eliminate grouping sets.
3. You used the exact same queries to "warm up" the data warehouse, and to test the performance. Some data warehouses (notably Redshift) aggressively cache intermediate compilation results, so they are much faster the second time they see a query or even a fragment of a query. To model a real user submitting queries interactively, you should use warmup queries that are similar to but not the same as the ones you use to measure performance.
4. You can solve the "vendor benchmarking their own product" problem by submitting a PR to our repo [4], which currently tests Redshift, Snowflake, BigQuery, Azure SQL DW, and Presto. We'd be happy to review it and endorse the timing if it meets our standards for fairness!
[1] https://fivetran.com/blog/warehouse-benchmark
[2] https://www.youtube.com/watch?v=XpaN-PqSczM
We had to figure out how to get something to compare to and understand how we were doing. We looked at the recent Giga Om Microsoft 30T tpch results [1]. They used sum of query times to compare results (sum of times in terms of fraction of an hour to run them all * cost per hour). I decided to do the same sum of query time, except with TPC-DS 10T. As TPC-DS official results provide the power run time, we can get query times from there.
We wanted to push ourselves to test 10TB TPC-DS. It was much more data, much larger intermediate results.
Some databases don't support grouping sets, and that means they can't run the official queries as you said. When we added support for them in MemSQL we made a choice to implement features that customers want; customers want to run TPC-DS without changing the queries, too for testing and other reasons. There are suggested rewrites for systems that don't have grouping sets. MemSQL didn't always have that feature. It's another way to distinguish system capability, I don't know why some of those other companies choose not to add support. I don't think using grouping sets or rewriting matters too much for our perf.
We never cache results; for us the warmup time separates the time to compile and code gen. Some companies do cache final or intermediate results.
I will update our blog post with some of this info, but it might not be completed until tomorrow. Since you were interested in the first run, aka warmup only time I will provide it here first. It's about a 5% difference.
MemSQL Sum of query time (avg of runs not including the warm-up time, already in the blog post): 6,494.74
MemSQL Warmup time only (just the first run, so includes optimize and compile time): 6,860.68
I work at MemSQL.
[1] https://gigaom.com/report/data-warehouse-cloud-benchmark/ [2] https://fivetran.com/blog/warehouse-benchmark
> How come you haven't go to 10T?
We don't see a lot of 10 TB single datasets "in the wild". We think 100 GB - 1 TB is more representative of the average real-world data warehouse user.
> for us the warmup time separates the time to compile and code gen
You should include compile time---your compile time is really good! Redshift sometimes takes longer to compile the query than to run it, though I think they may have improved this in the last 6 months.
1. Their database is run in asynchronous durability mode.
2. They specifically do the one thing that TPC-C says you shouldn't do, which is get really high throughput on a small dataset. TPC-C enforces that you scale your data-stored with the query throughput. CockroachDB maxes out at ~12.8tpmC/warehouse because its waiting at the legal maximum throughput, as opposed to running up the numbers in a way that's against the rules (and spirit) of the benchmark.
3. They make all the TPC-DS mistakes that georgewfraser points out elsewhere in this thread.
4. They run in read committed mode (they don't support anything higher), CockroachDB runs in serializable mode.
I ended up ranting about this on Twitter, so rather than reproducing everything here, I'm going to link to my rant there. Apologies for the cross-posting across fora: https://twitter.com/narayanarjun/status/1128393193941274624
1. MemSQL is running with synchronous replication in all these benchmarks. All data is stored on a 2nd machine before any transaction is acknowledged as committed. You’re right this is not as strong as running with both synchronous writes to disk and over the network. MemSQL supports this as well and results in about a 40 to 50% performance hit depending on the disk speed. Very few of our customers run in this configuration so we didn’t include it (the edge case of multiple machines losing power is not worth the performance hit for them).
2. Can you point me to what you’re describing in the TPC-C specification? I have never heard of what you’re claiming. TPC-C has maximum allowed latency requirements for the 5 transaction types it runs and also requirements around the mix of those transactions in the workload. The goal of the benchmark is still to run as many "New Order" transaction per minute while maintaining the latency requirements of the other unmeasured transactions running in the background (this is what tpmC stands for). We used the Percona TPC-C driver for MySQL to handle this (with a few small bug fixes).
3. The main thing we wanted to show is that our performance on TPC-DS is similar (better at some scale factors, slower on others) to data warehouses that specialize in running these types of queries. We likely should have provided more details (per query break downs and what not).
4. We used the Percona MySQL TPC-C driver with some changes to make the initial data loading faster. That driver uses the “FOR UPDATE” clause in MySQL instead of running in serializable isolation level.
I know you did a lot of work on CockroachDB. The point of the blog post was not to attack cockroach (I personally didn’t want to mention it at all), but to show how MemSQL is different. We are one of the few distributed SQL databases with competitive results on all 3 major TPC benchmarks.
2. What you're looking for is the 'Think Time' mentioned in the TPC-C spec[1] (table in 5.2.5.7). From 5.2.5.2, I quote:
> for each transaction type, the Keying Time is constant and must be a minimum of 18 seconds for New- Order, 3 seconds for Payment, and 2 seconds each for Order-Status, Delivery, and Stock-Level.
Chapter 4 is pretty thorough on elaborating on this. The comment under section 4.1.3 explicitly states:
> Comment: The maximum throughput is achieved with infinitely fast transactions resulting in a null response time and minimum required wait times. The intent of this clause is to prevent reporting a throughput that exceeds this maximum, which is computed to be 12.86 tpmC per warehouse.
Again, CockroachDB numbers are right up against this limit - because the database is waiting, as required! It's within ~99% of the maximum allowed. No bar is allowed to go more than 1% higher! So stacking a bar chart next to it that goes 10x higher is pretty misleading.
3. I'm pretty impressed that you can run all the TPC-DS queries. That's pretty impressive. But performance wise, there really isn't enough fleshed out, and given that the TPC-DS authors explicitly disavow the single metric that you use (power test numbers), is simply too little to claim parity to existing databases. That said, in this conversation I'm an OLTP guy; I'll let others more experienced with Data Warehouse benchmarking take this up, e.g. [3]
4. This one I'll concede that you are doing the appropriate thing as per spec (SELECT FOR UPDATE ensures serializability), but it's the single part of the spec that's not held up over time - the paper "Making Snapshot Isolation Serializable" is a great explanation of just what lengths you have to go to to prove that a set of transactions only provide serializable histories when run in a degraded isolation mode. That said, fair enough, no anomalies will be present due to Alan Fekete's proof. But do note that CockroachDB is doing a lot of extra work (work that MemSQL can elide, since it's simply not checking for isolation anomalies) to ensure that histories are always serializable[4].
5. While I don't work there, I did a lot of work specifically on benchmarking CockroachDB, and would like to politely request that you take down those bars for CockroachDB, since you're taking numbers that are shackled to the THINK TIME maximum and comparing them to a system that is not.
[1]: http://www.tpc.org/tpc_documents_current_versions/pdf/tpc-c_...
[2]: https://dl.acm.org/citation.cfm?id=1071615
[3]: https://twitter.com/gregrahn/status/1128448156180422656
[4]: I'll shamelessly plug my blog post on this for the reader interested in more about transaction isolation levels: https://ristret.com/s/f643zk/history_transaction_histories
Again, our goal here is not to have some showdown with cockroach. We don't really compete with each other. Our goal is to show the breadth of workloads MemSQL can run (fast in-memory point queries as well complex OLAP queries over large tables). None the less, we should have caught this before we published the article. I appreciate the correction.
If CockroachDB is concerned about THINK TIME enough to ask for the numbers to be removed, this would be a great opportunity for them to remove that limit and see exactly how much they could push the benchmark.
TPC-C requires that you increase the amount of "live data" if you want to display/advertise more performance. That's the benchmark's rule.
If you want to benchmark something else, that's fine, but then
1) don't call it "TPC-C" 2) don't compare with databases that play by the rules.
I would've though it'd be easier to write on HN (and I know for a fact that it's painful for me to read on Twitter).
This is a genuine question by the way. I followed the link, then couldn't be bothered to try and make sense of it and got to wondering why you would do that.
You're right, MemSQL doesn't support foreign keys as of yet, but none of these benchmarks require foreign key support. Two of them (TPC-H and TPC-DS) are a set of complex SELECT queries where foreign keys are not relevant at all. TPC-C is a write heavy benchmark, but the specification doesn't require foreign keys to be maintained (the data model does indicate the foreign key relationships though)[1].
These are unofficial benchmark results (not independently verified by TPC), so our interpretation of the specs may not be 100% correct, but I think we got it right as far as foreign keys are concerned.
[1] http://www.tpc.org/tpc_documents_current_versions/pdf/tpc-c_...
The only other OLAP database that I’ve used is Amazon’s Redshift and FKs are for “informational purposes”.
Or is it considered a transactional database or an analytics database?
[1] https://www.memsql.com/blog/the-need-for-operational-analyti...
I'm assuming that MemSQL works fine with that sort of configuration, rather than requiring you to lock your data up in some proprietary format.
Also, unlike the bad old days of on-premise platforms, you can try things out to see how they work. You could even do that with a public dataset first, to see how it works (see https://registry.opendata.aws/ for a list of these).
For example, there is an Amazon Customer Review dataset of over 160 million customer reviews - you could use that and try MemSQL for various use cases, then look at alternatives.
Disclaimer - clearly as a Kognitio employee I'd suggest you looked at us for analytics use cases, and you can see an example of sentiment analysis as scale using the Amazon Customer Review data set at https://kognitio.com/blog/sentiment-analysis-amazon-reviews-.... Also, a couple of articles on LinkedIn at https://www.linkedin.com/pulse/100-shades-grey-other-amazon-... for another piece of work on that same data, and https://www.linkedin.com/pulse/media-brexit-story-so-far-may... for a view on Global Media coverage of Brexit over time.
the past few years at memsql we have watched quietly as competitors have released numerous benchmarks. often one particular benchmark that coincidentally fits into a niche strength. often cleverly ignoring its weaknesses. often against an antiquated version of our software.
we are quite proud to be the only database company to produce such a varied array of impressive benchmark results. no niche benchmarks. no old versions. as close to the truth as it gets.
other databases in the mirror may be further behind than they appear
We didn't intend the benchmarks to be a sales pitch. We are proud of the performance of MemSQL and what it has been able to achieve in our customers' workloads. We wanted a way to show what we are capable of with concrete numbers. We chose the standard benchmarks because they are well understood, not because they were necessarily representative of any given customer.
In general, benchmarks are useful to understand the strengths and weaknesses of a product at a basic level and how it compares to its peers, but we strongly encourage anyone evaluating their options to do a proper POC comparison on their actual use cases.
Is the correlation 1:1? No. Is it still relevant when making a decision about what technology to use? Absolutely.
That's the idea of benchmarks, but it's not actually true of real benchmarks, because invariably some factors which can improve a particular benchmarks beaean inverse relation to performance for some other use cases of the technology.
I'm not having a problem.
> is that everybody already knows that,
Certainly, some people act like they don't, at least in the context of specific benchmarks.
And some people outright claim the opposite of what you say everyone understands, e.g., by claiming that benchmarks inherently correlate with every possible use of a technology.
> and plenty of folks can still extract value out of a benchmark anyway.
Understanding both the general issue and, ideally, the specific areas of potential concern in relation to particular benchmarks and your intended use is a big part of being able to effectively extract value from a benchmark.and, yes, lots of people do recognize those facts and extract value from benchmarks.
Others don't, and still apply benchmarks in decision-making, but it's less clear that they are extracting value. Confusing the measurement most readily available with the measurement most relevant to need is a common problem (and not just with benchmarks.)
You're throwing up your hands and saying, "TOO DIFFERENT FROM REALITY!"
Other folks don't do that, and while not 1:1, they are able to correlate the performance of a benchmark with the performance of their own use case.
Try not to get wrapped around the axle on "everyone", by the way, it's not literal.
Again, I'm not having a problem.
> is you can't figure out how to use benchmarks to understand how a system works
No, I have no partucular problem evaluating whether a benchmark is useful to a decision and if so, how.
Nor have I said anything indicating any such problem.
> You're throwing up your hands and saying, "TOO DIFFERENT FROM REALITY!"
No, I'm saying he naive statement upthread that benchmark performance correlates with every possible use of a technology is nonsense.
That's it.
> Other folks don't do that
Actually, some do, but that's neither here nor there.
> and while not 1:1, they are able to correlate the performance of a benchmark with the performance of their own use case.
Once again, yes, lots of people have the skill to figure out whether and in what way particular benchmarks have utility for their usecases. I explicitly said that in the post you are responding to.
That's very different than your claim that I reacted against, which is that any benchmark inherently correlates with every use case, which is—again—complete nonsense.
Oh. That's what you thought I said? Right. I didn't mean to say that, and if somehow I did say that I withdraw. Benchmarks are valuable if you know how to use them, but of course I agree with you, a benchmark isn't always relevant to every use case.
Side note, quoting many small portions of a person's comment and replying exclusively to the quoted bits is inferior to replying to them with full sentences/paragraphs. I've only ever seen what you're doing done by folks who are super interested in arguing and completely uninterested in having a conversation.
It definitely comes across like you do have a problem, which you have repeatedly stated you do not (and I believe you)! It just muddies the water, what you're doing here.
You can correlate lots of things. Poor countries increase penis size. Ice cream leads to murder. Cheese kills people by tangling them in their bedsheets. These things are strongly correlated. But they have no causative relationship.
To put it another way, a single benchmark test used find a solution is scientifically equivalent to finding one "average" man, making him run a series of athletic tests, and making his performance the benchmark for men's athletics.
Like I've been saying, benchmark results are plenty useful if you know how to use them.