1. It’s a simple GROUP BY on a single table. You’re basically just measuring the scan speed. Real queries are dominated by shuffles and the probe side of joins; these aren’t even present in this benchmark.
2. He runs the query repeatedly and takes the fastest time. This is far too cache-friendly. In this example, the intermediate stages of the query or even the result are probably just sitting in memory on the nodes after the first couple runs.
If you want to measure the performance of a data warehouse, you need to use more complex queries and not run the exact same query repeatedly.
edit: Coincidentally, I am giving a talk about data warehouse benchmarking TONIGHT in NYC. If you’re in NY and interested in this subject, please come! https://www.meetup.com/mysqlnyc/