Gorilla: A Fast, Scalable, In-Memory Time Series Database [pdf]
vldb.org
vldb.org
1) What/where exactly are they using GlusterFS for? Has Gluster fixed their scaling problems yet? Specifically the issue where new storage spaces/nodes were only available to new directories and files, but not existing directories? Granted, the last time I looked at this was 2009 or so, but it was a flaw due to their "no master node" topology.
2) FB has an entire team to manage Hadoop/HBase. This shows just how much of a beast that stack is. Anyone who has run Hadoop on "Internet time" knows what I'm talking about. It's great at running time insensitive, deferred compute jobs in an academic or scientific setting. It's really hard to keep it all 100% running in an on-demand setting. Aside, I couldn't imagine just working on 1 product in an operations setting as my full-time job. Boredom/fatigue must be a problem on that team.
3) I'd like to see more information on the networking side. What transport protocol? How large are the average updates in frame size? Etc etc.
We've built something similar to Gorilla in-house, so I'm happy to see that we've come to some of the same conclusions.
It's pretty easy to build a 2PB storage system on Ceph that the average group of sysadmins can run.
But I suspect that even the 32-bit kdb+ is going to be significantly faster than this gorilla.
I've used Matlab ages ago (before they used to JIT), so I can't compare, but I did have a chance to compare K and "straight" Numpy, and for my uses K won by a nontrivial margin (30-40% IIRC) despite using essentially the same solution, and Numpy having comparable implementation -- I guess that's where the cache and memory speed effects come in.
I did venture into using numexpr, which got Numpy a little faster than K, and I would assume Numba and PyPy (neither of which was in existence or usable at the time) might have also helped -- but plain Numpy vs. K was a clear win for K.
It obviously depends on the workload - but I specifically commented about Gorilla - which is a write-a-lot, read-a-little kind of workload, in which K excels and Numpy (as I remember it from 6 years ago) not so much.
If the underlying data type is 64 bit double, aren't they losing precision for integers greater than 2^53?
P.S. Their choice of venues is nice.