TileDB: Storing massive dense and sparse multi-dimensional array data
tiledb.io
tiledb.io
And TileDB, the company, formed out of it, has recently received funding [2] from Intel in Intel's latest tranche of investments.
[1] https://people.csail.mit.edu/stavrosp/papers/vldb2017/VLDB17...
[2] https://techcrunch.com/2017/10/19/data-is-the-name-of-the-ga...
It seems like the days we were stuck in particular or limited ways of thinking about databases/persistent storage are finally well and truly behind us! Now we have many awesome tools to choose from, better to have more tools in the belt than less.
Also in many time-series applications involving sensors whose readings don't fluctuate that much, process historians often apply deadband compression (i.e. store only one value if it is within a certain band). The type of compression is lossy and sometimes a bit controversial for high-fidelity uses, but often results in efficient storage.
[1] zfp, fpzip: https://computation.llnl.gov/projects/floating-point-compres...
There is another thing that is worth considering and that is the algorithms (and even the theory) that works well for compression of discrete sources are not well suited for compressing real numbers (floating point numbers aren't, but, they are the poor man's reals). On the theory side, this bothered Claude Shannon enough that he decided to revisit this later in his career to create rate distortion theory, he knew that there was some unfinished business in information theory.
We do have sort of a chicken and egg problem here, especially when we want store a lot floating points for a ML workload. Learning how to compress and learning the underlying distribution are equivalent problems. If we have already learned the model, then yes we could compress the data well. But when we haven't, then by definition we wouldn't have the knowledge to do a good job of storing the data in a well compressed form. After we have acquired the knowledge to compress well, we don't really need the compressed data anymore to learn the model, we already have it. One way to address this would be to do both incrementally and simultaneously.
Another cool research application of TileDB that extends the storage library with the VP9 codec can be seen here: https://homes.cs.washington.edu/~magda/papers/haynes-sigmod1...
Mentioning Matlab and Excel immediately puts the product in the category "they know what they are doing" as opposed to "another group of sophomores trying to reinvent data science".
I'm still waiting for a raw data dump with "avoid copies at all cost" access for very raw, very verbose vehicle and manufacturing data that is in 99.9 per cent never accessed, but must be analyzed when errors are detected late in the process. I.e. there's practically no transformation to apply during storage, but upon access, transformation must be done.
If TileDB is kind of like a more structured, more low-level struct oriented Redis, it is a very welcome addition.
I say this as someone who helped in small ways to develop open source alternatives to Matlab.
BTW I don't begrudge that TileDB has MATLAB support.
I see plenty of modern scientific/engineering workload in my day job. From what I see around myself it is usually a handful of people set in their ways that are holding others back by keeping a ridiculously overpriced tool alive that has comparable if not better alternatives. I gather from comments that it is different in Europe.
You are right that engineers are conservative which no doubt plays some role. But I'd say technical debt and legacy code plays much more of a role.
Most Engineers (Chem, Mech, Matls etc) are not exposed to code during university. Maybe it is changing now but it is slow process. Often first exposure comes when the engineer enters industry and is asked to work on existing model - usually under supervision of a senior engineer. You learn whatever the language the senior engineer knows and that is typically what they learnt from similar mentorship - it is often Fortran or C/C++. Our Industrial Process has not changed dramatically in last 30 years. So once efficiently written and accurate simulation code exists very little reason to rewrite (if given the choice to rewrite an existing model in whatever the current language du jour is or continue to hack on something that already exists most engineers - at least the ones I know, would chose to hack on existing code). For pretty much this reason there is a heap of Fortan that is still alive in my org (with roots that can be traced back to the '80's). I think it's probably worth mentioning a lot of engineers don't approach code thinking about algorithms - it's equations i.e. I need to write something to solve Bernoulli's equation - or Ergun's equation or similar. Linear programming (i.e solving simultaneous equations) is the other main reason to write code and hey MATLAB does this pretty well...
The other reason for being "gunshy" about new technology at least in my org is that a lot of the senior engineering people still remember when we got bitten by investing in Microsoft "Stack" in the late 90's. There were a lot of Modelling done with VB6 and Access which tuned out to be a technology dead end. Access databases in particular have plagued our org in many cases multiyear effort to migrate something out of Access. Open source probably protects against this but everyone wants to be sure any new technology will still be around in 15 years time. I think that is why people in my org are finally starting to look at R it has passed that initial hurdle and there is some confidence it is not just a fad...
Any chance I could pick your brains about using either Postgres or TileDB?
Thanks!
ingesting data into postgres makes sense for sparce data but not for dense data because it is waisting a lot of of space due to storage of coordinates with every data point and every weather variable. If you are using NOOA grib2 forecast files those are dense . Not to mention losing compression in postgres. TileDB will store data compressed, the dimension coordinates themselves will be compressed, plus column storage (each NetCDF variable) will make retrieval of dense weather data blazingly fast as oppose to postgres where you will have to scan the whole table
SciDB --- license: Affero ,software size: 5GB ,data model: ACID ,focus on: dense ,dimensions: integer ------
TileDB license MIT ,software size <1MB , data model: eventual consistency via fragments , focus on: dense, sparse ,dimensions are integer, floats
It doesn't seem like the best application. On one of their pages it mentions they ingest BAM records, which is for biological sequences. I'm guessing some DNA storage applications.
This brings up a question: in what fields does one find heavy use of large, sparse matrices that need to be persisted and queryable?
In my mind, sparse matrices typically occur in the context of graphs/relationships, e.g. PageRank, logistic networks, adjacency matrices, etc. They also tend to be a property of Hessian matrices (2nd order derivative for a multivariate system). But typically these are intermediate quantities that are discarded after a computation completes.
TileDB supports both dense and sparse arrays. It was designed around the concept of handling sparse arrays but dense arrays can be thought of a degenerate case of sparse array storage in TileDB. For dense arrays tile extents are contiguous and we don't materialize the coordinate values. This way all the concepts are the same and we can capture both use-cases. Sparse annotations to dense array values, such as NA or Null handling can also be captured as a sparse array fragment layered over a backing dense array.
I agree with you that for most use cases, storage will be dense. But it is useful to have one system that can handle both representations efficiently, with the sparse case not added on as an afterthought (it also makes the system simpler).
As an ML guy I can say that sparse is pretty common in ML. Text data, market basket data, graph Laplacians, adjacency matrix, sequence fragment data are all rich in sparse matrix computation and storage operations. For a moment I thought the "not" was a typo. Lack of support for sparse matrices often becomes a serious inconvenience in a tool, very happy that TileDB folks have given thought to the sparsity requirement from get-go
A general comment: TileDB’s vision goes beyond that of the HDF5 (or any scientific) format. Considering though the quantities of HDF5 data out there (and the fact that we like the software), we are thinking about building some integration with HDF5 (and NetCDF). For instance, you may be able to create a TileDB array by “pointing” to an HDF5 dataset, without unnecessarily ingesting the HDF5 files but still enjoying the TileDB API and extra features.
Also saw in your documentation that you're concentrating on lossless compression right now, which makes complete sense. However, as a scientist, I just want to put a vote in for lossy compression too: it's not uncommon to work with large datasets given in float64 (because float64 is used for any intermediate processing steps), but that actual final precision we need to store is much less than that, but we're stuck with these huge binary files.