JuliaDB
juliadata.github.io
juliadata.github.io
As someone who's in the midst of a plan to migrate off of Apache Spark, because we've discovered that, for our particular purposes, a single-machine implementation in a language with better crunching and multithreading chops than Java, can easily achieve comparable performance to a Spark cluster, while being good deal easier to develop and maintain, this comment strikes me as being very poignant.
I have come to suspect that the big data ecosystem is a castle that was largely built on a nice, thick technical foundation that is largely composed of pointer chasing.
The amount of money and engineering resources thrown at a problem that actually has a pretty simple solution (integrate your db and language, make your program and language runtime fit in L1 i-cache, optimize the hell out of your language) is a little disheartening. But hey, it keeps people employed managing and debugging an unholy mess of DBs, K8 clusters, load balancers, KV stores, and message brokers..
The language itself is extremely minimal, so the learning curve is actually much lower than most languages (it forces you to write in a "vectorized" style, which is not too different than how numpy/R/matlab works). I was able to get up to speed in a couple of days.
Source: Used to work with kdb+ (still do sometimes). Not a shill (I don't even like k/j/q/APL).
The APL family takes some getting used to, as it's a significantly different paradigm from Julia and Python looking languages: arguably as different a paradigm as Lisp. It's also fairly natural if you've done any work with Matlab, NumPy or other array oriented packages, but you do have to put in the work.
It doesn't get as much attention here, but Jd is pretty good also. It's also much lower barrier to entry than Kx if you're building a personal project with it (aka its free for personal use). J is a more complex language than K, but I like it better; more oriented towards mathematics.
Its not a trivial language to learn, you have to be ok with very terse languages in general, and ok using the language as a way to mentally model the algorithm. Somewhat "mathematically".
I'm not using/working with it any more, haven't in a few years. Most of my recent work is in Julia, by choice.
Two very different philosophies at work here. Julia provides a sane, more or less classical language to work with, inclusive of numerous macros and modules to help drive productivity. k/q provides a fairly modern apl like experience, leveraging Iverson's idea of a programming language being a "tool for thought".
FWIW, the forces behind k/q (Arthur Whitney, et al) are now working on Shakti[1]. Looks like an interesting project, more of a platform than kdb was.
Bringing this back to JuliaDB, I had been hoping for a very julia centric/idiomatic persistence engine as well as an analytical engine. I'm no fan of SQL or its variations. In reading this, I've learned DataFrames is supposed to be this. However, my impression working with DataFrames has been that it is woefully slow on small data sets (100-ish rows). It cannot handle one of the slightly larger 50k line tabular data sets that I am importing via excel sheets.
For that work, I may need to resort to Perl (my fall back programming/data munging language). This is important, as Perl has a capability of tieing (binding) a variable to an underlying data structure/file/... . This enables trivial persistence, in a completely idiomatic manner (update/insert into the tied variable by setting a value). It would be interesting if this capability existed within julia as well.
[2] https://kx.com/media/2018/03/joe_preview.jpeg (actual image of me talking)
Because right now I'm using a DataFrame that has 7m rows and 50 columns and it works very well.
I should look into this more. Thanks for letting me know it can handle things that size.
Also you can set up clusters using a very nice IPC system. You can seamlessly just send over code and data with very little fuss.
Most the programming platforms and libraries that one might choose instead of Hadoop & friends didn't exist back then, or weren't or weren't particularly robust. So there weren't a lot of games other than Java in town. And Java (idiomatic Java, not benchmark-friendly Java) tends to reduce the volume of work you can squeeze out of a given hardware setup by introducing a hefty does of memory overhead.
Really fun reading on exactly this issue - definitely changed my view on Spark and compute clusters. People definitely don't realise just how powerful a modern computer is, and how far you can get on a single one.
You should take a look at genomic data. The volume of data is one of the main challenges[1].
[1] https://spectrum.ieee.org/computing/software/the-desperate-q...
e.g. Cluster of scalable ARM SBC nodes or even multi-architectural nodes could be highly efficient.
I've been exploring this kind of setup for a while now, frameworks like Dask, Ray[1], Modin can do this with Python to some extent. But they are still finicky and Dask seemed more stable than other frameworks for this setup.
I wanted to try out language level distributed computing setup, but Julia required same setup (OS/Arch/Path) replicated on all their nodes last time I visited it.
I feel, distributed computing as such hasn't got much love as it deserves in the consumer market, especially since many have several computing units in their houses now(PC, tablet, smartphone, Watch, TV, Game console). If interoperable distributed computing layer was fundamentally baked in with all modern operating systems minimising the latency with Network, Storage, Memory; Then we could leverage huge compute power on demand, which not possible without investing lot money in single compute unit nowadays.
Then again, Compute power is a major strategic advantage for the manufacturers. Why would Apple share its A13X for compute with a Snapdragon or Intel?
[1]https://gist.github.com/heavyinfo/aa0bf2feb02aedb3b38eef203b...
For if you can fit it in RAM on a high-end machine, and on a single harddrive in a low-end machine.
and its practically a more useful case. Since if data is truely big many algorthms can't be used (since if it is O(n^2) its not practical).
See diskframe.com a single machine focused framework
More natural is extending Tables.jl (like DataFrames and JuliaDB does). and we continue to build more tools that are table agnostic, and have APIs that DataFrames and a distributed table package special case when they can do it more efficienctly
I'm working on a set of functions to work with many DataFrames (or anything) in parallel and do it out of core if possible, it's basically like JuliaDB but with a FileTree abstraction rather than a table abstraction.
Their README says:
The package currently provides working implementations for in-memory data sources, but will eventually be able to translate queries into e.g. SQL. There is a prototype implementation of such a "query provider" for SQLite in the package, but it is experimental at this point and only works for a very small subset of queries.
Still early days, but sounds like they're working on the same problem.
I wish the Julia ecosystem was a little more integrated: there are a lot of different competing libraries that ostensibly do the same thing. Python has the advantage that it's obvious what you should use: numpy, Pandas, scipy, statsmodels, matplotlib, etc.
With Julia, it's less clear. Though I think part of the reason is that actually releasing a new scientific computing Python library is incredibly difficult and requires a lot of expertise.
Julia makes it pretty trivial for anyone to contribute a model that has excellent performance. This fragmentation is a common problem among expressive languages.
It also uses strongly typed tables (e.g. Table<int, string> etc), whereas DataFrames is loosely typed. Again I think this is a good decision (though it does grate with the Julia JIT's property of "being slow" the first time you run a function on a new type)
Finally its split into IndexedTables and NDSparse is again a good design decision that I have not seen replicated in any other dataframe library.
It just seems all around better designed.
On the other hand it is verging on being unmaintained.
So I think it all mostly has to do with the fact that Julia is much newer and many even foundational areas are still in a state of flux. In a few decades the dust will have mostly settled. Of course, in the areas where people are pushing the envelope, it's good that a lot of approaches are tried out before the world settles on a winner.
But nonsustsined efforts will fall by the way side and true gems will emerge as the clear front runner like DataFrames.jl
The Lisp curse?
1) don't put "julia" in the package name
2) don't use abbreviations in the package name
Of course, JuliaDB is probably older than these conventions, but it's amusing nonetheless.
If it's a full database engine, it might be usable from other languages. So the name makes sense.
It's cool that it's pure Julia, so it's instantly portable everywhere Julia runs, and the code is safer than it could be were portions of it written in C.
This thing is pure Julia.
This being any faster than Pandas is a huge compliment to Julia the language in general and its JIT compiler in particular.
Indeed. This 'whole stack under the same language' is an important feature to have. You get to reap the advantages of an improved JIT. End to end autodiff is easier. Some of these things are a problem, for example, in PyPy because of the Python <-> C bridge.
But it's called a db
In fact, I've been using Julia for work and following the ecosystem since version 0.4 (we're at 1.5 now), and I'm still not sure what JuliaDB is. No doubt this is mostly due to me not having reason to look very deeply (and/or not being very perceptive), but certainly doesn't feel like the marketing copy is giving me any help...
> JuliaDB is a pure Julia analytical database. It makes loading large datasets and playing with them easy and fast. JuliaDB needs to support a number of features: releational database operations, quickly parsing text files, parallel computing, data storage and compression.
Got this from https://juliadb.org/talk/juliacon2018shashi/ which appears to go into more details
> This talk is a bottom-up look at the construction of JuliaDB. We will talk about the scope and implementation of underlying building block packages, namely IndexedTables, TextParse, Dagger, OnlineStats and PooledArrays.
It doesn't seem to store data in any way - so it is definitely a data processing engine, but a database without an "INSERT" command feels a little off.
From that point of view, this looks a lot like what original MapReduce did - the data lives outside it & is referenced as urls, but the engine itself does processing out-of-core and in-memory for very large datasets.
I agree that it's a bit odd to not have a direct analog of SQL's INSERT, but you can definitely add rows to an existing table by making a new one and doing a merge operation.
Side note: this is what we tell every startup about talking to HN: Lead with a clear statement of what your company does. If you don't, the discussion will consist of "I can't tell what your company does". Same for open source projects of course.
JuliaDB is more like Dask, or DataFrames.jl - it provides that computation functionality.