Array Databases: Concepts, Standards, Implementations [pdf]
rd-alliance.org
rd-alliance.org
My two cents (as the author of Xarray, one of the Python libraries mentioned in this report) is that it’s questionable whether we need “array databases” at all. Certainly we need to be able to store arrays and compute with them, but do we need an integrated solution that does both at the same time with a query language that looks like SQL? Maybe not, in an era of cloud computing, prolific open source software and when everyone who works with big array datasets already knows Python.
This query language that looks like SQL is an official part of SQL now [1]. Surely there is place for integrated DB solutions that let you work with both relational and array data in one place? There are more benefits in this than just performance/scalability. Think of building services on top of big array datasets, beyond one-off data science experiments.
In kdb you go from primitives to arrays to tables. In SQL you go from primitives straight to tables which makes it cumbersome to do any simple one column or array ops. Such as excluding a column from a select expression.
Compare first class array support in sql vs hypothetical programmable sql
select table.* - {name, age}
from table
vs select (
select column
from table
where column
not in ('name', 'age')
)
from tableThis make a lot of sense. Because primitives ARE "tables", columns ARE "tables".
A primitive is a relation of one column/row. This is what allow you to do:
SELECT * FROM (SELECT 1)a
What sql/rdbms not do it well is to exploit this very well.Many orgs these days store all data in data lake shared-disk architectures and pull down the subsets. The performance hit of pulling down data over high bandwidth channel such as s3 - ec2 is much more reasonable to companies than storing everything on expensive compute instances just so that the "data would be there" ready for querying if somebody ever needs it.
That said: you may call it "suspicious" if the authors come to that conclusion, but on the same grounds it is likewise suspicious if the writer of a tool that has not excelled doubts the results :)
Let us rather look at facts - the benchmark is published and open, and actually similar figures have been reported by other, completely independent benchmarks. The paper has undergone tough scrutiny by 5 independent experts in the field before publication.
Doubting about the value of databases reminds me of the old times of COBOL vs SQL: "we don't need SQL data management". Incidentally, IT world since then has embraced databases of all kinds...and exactly arrays should be the big exception? That does not make sense. Tools will need to accept that there are other tools as well, and ideally discussion is merit-based.
"everyone who works with big array datasets already knows Python"...well, if the only tool you know is a hammer then...ya know. There are so many more worlds than just your comfort zone, python! Just think of R, for example.
PS: I am not questioning xarray - in projects we have used xarray as a frontend to rasdaman, and this combination works like a charm: python wrapper around scalability and federation, connected through the open OGC standards.
I agree, formulating open benchmarks is great work, and there's nothing suspicious about performing well on benchmarks you write. It's just worth keeping in mind: https://matthewrocklin.com/blog/work/2017/03/09/biased-bench...
My initial remarks were a little careless and overly provocative! I don't doubt that there are cases where a true "array database" provides value. I do think the use-cases are less clear for arrays than they are for tabular data, because the users of arrays tend to be more sophisticated.
Xarray definitely takes a different philosophical approach based on its roots in the Python data science ecosystem, compared to "all in one" solutions like a full array databases.
my bad, sorry!