When Google published the Dremel paper in 2010, it explained how this structure is preserved within column store.
...
The definition and repetition levels encoding is so efficient for semistructured data that other open source columnar formats, such as Parquet, also adopted this technique.
Parquet is an open-source reimplementation of the columnar storage format described in Google's 2010 Dremel paper. Capacitor is Google's next-generation columnar storage format.Why switch to what everyone else is using then?
(See http://tech.marksblogg.com/billion-nyc-taxi-rides-bigquery.h... vs all other benchmarks for the same dataset Mark got)
Disclaimer: I'm Felipe Hoffa and I work for Google (https://twitter.com/felipehoffa). But you can try BigQuery in the next 5 minutes and check the speed claims :).
> BigQuery is faster than anything else I've seen. > Why switch to what everyone else is using then?
This really depends on how much faster. For a marginal drop in performance, many would think it's worthwhile to stick with an established format. That said, I'm willing to believe the performance delta for BigQuery is worth it :)
We dump the output of spark jobs into BQ for exploration and having to produce JSON in addition to parquet is an irritating (and expensive) overhead.