Capacitor, BigQuery’s next-generation columnar storage format
cloud.google.com
cloud.google.com
The performance gain from columnar storage is the compression ratio. And by ordering similar attributes together, it is going to greatly reducing the entropy between rows, which in turn leads to high compression and better performance.
The trick is to smartly select the which column you are going to have all the rows sorted upon.
I my previous company, we are using Redshift to encode 1 billions rows, and the simple change to let the table sorted by user_id reduce the whole table size by 50%, that is half a TB of disk storage, the improvement is nothing more but jaw-dropping. I think Google here just takes this trick into a more systematic method, which is really neat.
To point out, in columnar storage system, take ordering into account. Try some ordering that you feel could maximize the redundancy between rows, usually it is going to be primary id that is most representative of the underlying data. You don't need to have a fancy system like this one to leverage this power idea, it could apply to all columnar systems.
Here's one paper: Sorting improves word-aligned bitmap indexes, http://arxiv.org/abs/0901.3751
I wonder if they could share more details on how this is handled.
You can do adaptive layout rewriting at either the page or shard level, depending on the design. There are advantages and disadvantages to both models. Some designs can do layout conversion in place without the need for garbage collection but it is much trickier to do correctly.
Why switch to what everyone else is using then?
(See http://tech.marksblogg.com/billion-nyc-taxi-rides-bigquery.h... vs all other benchmarks for the same dataset Mark got)
Disclaimer: I'm Felipe Hoffa and I work for Google (https://twitter.com/felipehoffa). But you can try BigQuery in the next 5 minutes and check the speed claims :).
> BigQuery is faster than anything else I've seen. > Why switch to what everyone else is using then?
This really depends on how much faster. For a marginal drop in performance, many would think it's worthwhile to stick with an established format. That said, I'm willing to believe the performance delta for BigQuery is worth it :)
When Google published the Dremel paper in 2010, it explained how this structure is preserved within column store.
...
The definition and repetition levels encoding is so efficient for semistructured data that other open source columnar formats, such as Parquet, also adopted this technique.
Parquet is an open-source reimplementation of the columnar storage format described in Google's 2010 Dremel paper. Capacitor is Google's next-generation columnar storage format.We dump the output of spark jobs into BQ for exploration and having to produce JSON in addition to parquet is an irritating (and expensive) overhead.