Metaprogramming Python For Big Data
tuulos.github.io
tuulos.github.io
I'm happy to answer any questions.
The data format is a collection of pre-aggregated row and column vectors, encoded with variable length integers and run-length encoding. I should give a separate presentation about this.
Please do. Also, any change of releasing this?
I know continuum was working on a version of memmapped numpy (blaze, iirc) arrays which looked really interesting.
Pandas uses NumPy internally. You could use Deliroll as a replacement for NumPy in Pandas to get a nice interactive environment for amounts of data that can't be easily handled with plain NumPy.
I passed the slides along to my team to see what they think. If they just impress upon the team that we need to store something other than just 0x0A delimited text files, I'll consider it a win.