sincere question - Is your only data ingest mechanism by uploading gzipped csvs, or other files? it seems that if people really have big data, then by definition that approach won't work
We'll be supporting cloud file bucket locations soon: S3, etc. We're also working on handling streaming data, e.g. logs.
Have signed up for the beta. Look forward to checking it out. I've looked at the Google prediction API, but it doesn't do what I need.
What do you need?
Most of my data is too big for CSVs but too small to justify distributed storage. I use HDF5 with chunking and column compression. I think many other people in the sciences and finance also do this (along with using NetCDF).