Anybody who comes up with a new "Big Data" solution has to realize that the 90% of the data users are vested deeply in Hadoop and the ecosystem around it. Having tools like PrestoDB, Hive, Shark on the top of HDFS makes it really hard to newcomers to convince companies to invest in something else. If data was a greenfield territory these project would be more viable. Btw. on this note, PrestoDB just getting the right features like join optimization, support for nested structures, etc.
Maybe in your case they have Hadoop, but not everyone is satisfied with Hadoop or wants to create new systems with it. Choice is good, as long as we can differentiate the choices by the values they provide.
we here at crate were living in the hadoop ecosystem for some time. but when we came across the beautiful architecture of netty.io (async, event-driven) and the way elasticsearch orchestrated it - since that time we know there will be room for newcomers :)
Well sorry my buzzword filter removed half of your sentence. :) I understand that engineers can get excited about async and event driven but these mean absolutely nothing to your endusers. I could implement your service with blocking IO and non-event driven code and still get the same performance. It simply does not matter that much. The key to Hadoop's success is scalability and predictable performance. Don't get me wrong, I think Hadoop is one of the worst ecosystems I have ever seen in my entire life, the code quality makes me cry sometimes, but putting all these aside, projects like PrestoDB making Hadoop viable and keeping it alive in the long run. Your project looks interesting but without extensive performance testing and proving that your TCO is lower than Hadoop's, and having the same features as Hive, PrestoDB does will be really hard to break in to this market. Again, I would be the happiest person to see something more sane than Hadoop on the market.
I like Hadoop tonnes so I would love to get an injection of reality and perspective: what are the insanities of Hadoop?
sorry for my ignorance, but Crate and "SQL on top of Hadoop" fulfil quite different needs. How would you handle tens of thousands concurrent queries on a Hadoop platform? Read and write at the same time? On the other hand side - crate will never be able to batch-oriented, complex map/reduce workload.
Well I guess not everyone can hire extra engineers to run and keep Hadoop happy so this is a viable solution for smaller teams as long as the performance is on par. I am no big data guy nor have I worked with Petabytes of data but as a front end focused developer running the tech part of my startup, I'd rather deploy Crate and worry less about it than pay tons of money for a DBaas solution.
totally; the main reason for me is the idea of substantial long term investment by large support communities (including me!)