Docs: https://docs.tiledb.com/developer/
Website: https://tiledb.com/
Earlier HN post: https://news.ycombinator.com/item?id=15547749
Disclosure: I am a member of the TileDB team.
Docs: https://docs.tiledb.com/developer/
Website: https://tiledb.com/
Earlier HN post: https://news.ycombinator.com/item?id=15547749
Disclosure: I am a member of the TileDB team.
The webpage does not tell me exactly what TileDB is, so I'm having a hard time getting understanding what's underneath all the marketing mumbo-jumbo.
But if it is indeed a better SQLite/Parquet, DAMN, I'm so going to spread the gospel of this to all my students.
For a lot of what I do, I want a hierarchical containment system- the equivalent of folders with files. And the files themselves are leaves in the hierarchy, containing multidimensional array data. WOrks great when the arrays are composed of fairly straightforward payloads, like float[x][y][z] but also works if your array values are structs. Much of the value in zarr and tiledb comes from specifically how they arrange the arrays, for convenient read access to slices of the arrays. Access is going to look like: ages = root["user"]["age"][100:100:2]
Parquet is mostly a column file format, but with nesting. I'd use it to store large amounts of structured data with a relatively straightforward schema, although the schema itself can be fairly nested so some records have very complex structure. Access would often be in a loop over all records: for record in records: if record.user.has_age(): print("User age:", record.user.age)
SQLIte is a library/CLI that implements a relational database. It has a SQL interface and stores data using classic relational DB approaches, including secondary indices, etc, and permitting joins directly within the engine: SELECT age FROM user WHERE user.country == 'Bulgaria'
The good news is that we offer efficient integrations with MariaDB, PrestoDB and Spark, so you can directly process SQL queries on TileDB data via those engines (which work even for dense data). With MariaDB, we even have a embedded version which allows running SQL queries directly from Python[2]. This combines the ease of use of sqlite with MariaDB's speed and TileDB's fast access to AWS S3 (and Azure Blob Store in the next version).
[1] https://docs.tiledb.com/main/use-cases/dataframes
[2] https://docs.tiledb.com/developer/api-usage/embedded-sql
Disclosure: I am a member of the TileDB team.
Taking a look at the compression filters, it seems that they are mostly suited for byte-oriented and integer data? The bit width reduction filter, for example, seems to act only on integers, and it doesn't look like any floating-point-specific compressors (like fpzip) are implemented -- just general-purpose ones like bzip2.