For the bioinformatics datasets I work with it is not cost efficient to load everything into a database. For some of these datasets (e.g. GWAS - Genome Wide Association Studies which are essentially sparse matrices) it might be interesting to explore with graph queries. I guess my ideal would be to have a GraphBLAS equivalent to Spark SQL queries working across files in cloud storage / NAS.