So indeed, it does not matter whether the data is stored on disk or in memory --- column-stores are a win for these types of workloads. However, the reason is totally different.
He says that for tables on disk, column layout is a win due to less data transferred from disk. But for tables in memory, the win is due to the fact that you can vectorize operations on adjacent vaules.
I also think the convergence of these two approaches is interesting: one is from a database POV; the other is from a programming language POV.
Column Databases came out of RDBMS research. They want to speed up SQL-like queries.
Apache Arrow / Pandas came at it from the point of view of R, which descended from S [1]. It's basically an algebra for dealing with scientific data in an interactive programming language. In this view, rows are observations and columns are variables.
It does seem to make sense for there to be two separate projects now, but I wonder if eventually they will converge. Probably one big difference is what write operations are supported, if any. R and I think Pandas do let you mutate a value in the middle of a column.
-----
On another note, I have been looking at how to implement data frames. I believe the idea came from S, which dates back to 1976, but I only know of a few open source implementations:
- R itself -- the code is known to be pretty old and slow, although rich with functionality.
- Hadley Wickham's dplyr (the tbl type is an enhanced than data.frame). C++.
- The data.table package in R. In C.
- Pandas. In a mix of Python, Cython, and C.
- Apache Arrow. In C++.
- Julia DataFrames.jl.
If anyone knows of others, I'm interested.