Druid: fast column-oriented distributed data store
druid.io
druid.io
Then most places will augment this with a range of database solutions depending on how structured/clean the data is and the various workloads. Cassandra, HBase, MongoDB, Teradata, Oracle, ElasticSearch, Greenplum are all pretty common place in most enterprises.
And of course the ridiculous array of SQL engines on top of HDFS/S3/whatever else e.g. Hive, Spark SQL, Presto, Drill, SAP etc.
I wrote my interpretation of the current open source data landscape here for anyone interested: http://imply.io/post/2015/11/04/big-data-zoo.html
"Big data" stream processing is obviously related as its a dataflow programming model, but it's still very different in practice. The streaming abstraction is generally more free-form, and not realized directly in hardware. I contrast the two kinds of streaming in Section 2 of a paper from a few years ago: http://www.scott-a-s.com/files/pact2012.pdf
- Harish.
The biggest problems with SparkSQL is simply in its limited support for ANSI SQL. It's getting better with every release but not nearly quick enough.
Can you be specific about ANSI SQL compliance requirements. Spark SQL is closing the gap on Hive SQL; both have decent support for analytical queries: Cubes/Rollups/Windowing etc. The only major gap between Spark and Hive SQL I know off is SubQuery predicates(exists/not exists).
When I was at Optimizely, my team chose Druid for a large-scale analytics application, after a pretty extensive benchmarking. It was very impressive, though not trivial to set up.
Blog post with a bit more detail: https://medium.com/engineers-optimizely/slicing-and-dicing-d...
> we were delighted to find an existing community cookbook, chef-druid, that can configure and deploy a druid cluster. However, that community cookbook hasn’t been updated for nearly a year and does not support the latest version of druid. We have therefore forked off our own version, optimizely/chef-druid, which supports the latest version of druid.
Why didn't you create a pull request to the upstream project? That way, all people who are using the original cookbook could have directly benefit from your improvements, without having to discover the new fork.
Speculating personally, it was probably just something that happened for speed/ease - as opposed to trying to get an apparently abandoned repo rolling again.
We did make optimizely/chef-druid open and public though, and while I'm not with Optimizely any more, I'm sure the good folk there would be more than happy to merge back into the original repo if anyone was still maintaining it and wanted that to happen.
That's a short description of 12 alternatives I wrote up. Druid does look interesting, it appears to have got the architecture pretty spot on and I'm going to look into it more.
Elasticsearch and Redshift are.
Seriously: What, specifically, do you disagree with, in the context of your use case, and what is your use-case?
If this transition is easy without reworking infrastructure, the solution is far more attractive.
http://www.postgresql.org/docs/current/static/functions-wind...
With that said, it works very well, but it definitely came at the cost of a good dose of sanity.
There's also a production-ready docker distribution: https://hub.docker.com/r/imply/imply/