As someone who's in the midst of a plan to migrate off of Apache Spark, because we've discovered that, for our particular purposes, a single-machine implementation in a language with better crunching and multithreading chops than Java, can easily achieve comparable performance to a Spark cluster, while being good deal easier to develop and maintain, this comment strikes me as being very poignant.
I have come to suspect that the big data ecosystem is a castle that was largely built on a nice, thick technical foundation that is largely composed of pointer chasing.
Most the programming platforms and libraries that one might choose instead of Hadoop & friends didn't exist back then, or weren't or weren't particularly robust. So there weren't a lot of games other than Java in town. And Java (idiomatic Java, not benchmark-friendly Java) tends to reduce the volume of work you can squeeze out of a given hardware setup by introducing a hefty does of memory overhead.
You should take a look at genomic data. The volume of data is one of the main challenges[1].
[1] https://spectrum.ieee.org/computing/software/the-desperate-q...
See diskframe.com a single machine focused framework
The amount of money and engineering resources thrown at a problem that actually has a pretty simple solution (integrate your db and language, make your program and language runtime fit in L1 i-cache, optimize the hell out of your language) is a little disheartening. But hey, it keeps people employed managing and debugging an unholy mess of DBs, K8 clusters, load balancers, KV stores, and message brokers..
The language itself is extremely minimal, so the learning curve is actually much lower than most languages (it forces you to write in a "vectorized" style, which is not too different than how numpy/R/matlab works). I was able to get up to speed in a couple of days.
Source: Used to work with kdb+ (still do sometimes). Not a shill (I don't even like k/j/q/APL).
Also you can set up clusters using a very nice IPC system. You can seamlessly just send over code and data with very little fuss.
Its not a trivial language to learn, you have to be ok with very terse languages in general, and ok using the language as a way to mentally model the algorithm. Somewhat "mathematically".
I'm not using/working with it any more, haven't in a few years. Most of my recent work is in Julia, by choice.
Two very different philosophies at work here. Julia provides a sane, more or less classical language to work with, inclusive of numerous macros and modules to help drive productivity. k/q provides a fairly modern apl like experience, leveraging Iverson's idea of a programming language being a "tool for thought".
FWIW, the forces behind k/q (Arthur Whitney, et al) are now working on Shakti[1]. Looks like an interesting project, more of a platform than kdb was.
Bringing this back to JuliaDB, I had been hoping for a very julia centric/idiomatic persistence engine as well as an analytical engine. I'm no fan of SQL or its variations. In reading this, I've learned DataFrames is supposed to be this. However, my impression working with DataFrames has been that it is woefully slow on small data sets (100-ish rows). It cannot handle one of the slightly larger 50k line tabular data sets that I am importing via excel sheets.
For that work, I may need to resort to Perl (my fall back programming/data munging language). This is important, as Perl has a capability of tieing (binding) a variable to an underlying data structure/file/... . This enables trivial persistence, in a completely idiomatic manner (update/insert into the tied variable by setting a value). It would be interesting if this capability existed within julia as well.
[2] https://kx.com/media/2018/03/joe_preview.jpeg (actual image of me talking)
Because right now I'm using a DataFrame that has 7m rows and 50 columns and it works very well.
I should look into this more. Thanks for letting me know it can handle things that size.
The APL family takes some getting used to, as it's a significantly different paradigm from Julia and Python looking languages: arguably as different a paradigm as Lisp. It's also fairly natural if you've done any work with Matlab, NumPy or other array oriented packages, but you do have to put in the work.
It doesn't get as much attention here, but Jd is pretty good also. It's also much lower barrier to entry than Kx if you're building a personal project with it (aka its free for personal use). J is a more complex language than K, but I like it better; more oriented towards mathematics.
Really fun reading on exactly this issue - definitely changed my view on Spark and compute clusters. People definitely don't realise just how powerful a modern computer is, and how far you can get on a single one.
e.g. Cluster of scalable ARM SBC nodes or even multi-architectural nodes could be highly efficient.
I've been exploring this kind of setup for a while now, frameworks like Dask, Ray[1], Modin can do this with Python to some extent. But they are still finicky and Dask seemed more stable than other frameworks for this setup.
I wanted to try out language level distributed computing setup, but Julia required same setup (OS/Arch/Path) replicated on all their nodes last time I visited it.
I feel, distributed computing as such hasn't got much love as it deserves in the consumer market, especially since many have several computing units in their houses now(PC, tablet, smartphone, Watch, TV, Game console). If interoperable distributed computing layer was fundamentally baked in with all modern operating systems minimising the latency with Network, Storage, Memory; Then we could leverage huge compute power on demand, which not possible without investing lot money in single compute unit nowadays.
Then again, Compute power is a major strategic advantage for the manufacturers. Why would Apple share its A13X for compute with a Snapdragon or Intel?
[1]https://gist.github.com/heavyinfo/aa0bf2feb02aedb3b38eef203b...
For if you can fit it in RAM on a high-end machine, and on a single harddrive in a low-end machine.
and its practically a more useful case. Since if data is truely big many algorthms can't be used (since if it is O(n^2) its not practical).
More natural is extending Tables.jl (like DataFrames and JuliaDB does). and we continue to build more tools that are table agnostic, and have APIs that DataFrames and a distributed table package special case when they can do it more efficienctly
I'm working on a set of functions to work with many DataFrames (or anything) in parallel and do it out of core if possible, it's basically like JuliaDB but with a FileTree abstraction rather than a table abstraction.
Their README says:
The package currently provides working implementations for in-memory data sources, but will eventually be able to translate queries into e.g. SQL. There is a prototype implementation of such a "query provider" for SQLite in the package, but it is experimental at this point and only works for a very small subset of queries.
Still early days, but sounds like they're working on the same problem.