No one ever got fired for using Hadoop on a cluster (2012) [pdf]
research.microsoft.com
research.microsoft.com
On a meta-note, I felt like the abstract wasn't very clear and I had to skim the rest of the doc to figure out what was going on.
http://www.frankmcsherry.org/graph/scalability/cost/2015/01/...
http://www.frankmcsherry.org/graph/scalability/cost/2015/02/...
No kidding. These arrangements are much more entertaining and informative.
In a rush to get to the next big thing, many have taken a lot of risk with Hadoop. Some with the resources to really support/make Hadoop their own have had a lot of success. But I think there are fewer who have success stories.
I just don't see a customer facing SLA sensitive workload fitting this well. At least I would not sleep well being oncall... :)
Generally, I agree though, in an extremely demanding SLA environment... I probably wouldn't sleep well in that case either!
I think that was kind of the point of the paper as presented - Hadoop is seen as a panacea, when in reality there might be other, simpler approaches that work just as well or better. It really does depend on the use case, the volume/types of data, cost/requirements, etc.
For that matter, what the Hadoop ecosystem "is" (vs. just the Apache Hadoop project itself) means so many things now. HDFS (storage), YARN (distributed job/resource management), mapreduces, Hive, HBase, etc. vs. new engines, like Apache Spark, for example, which can run inside or outside of Hadoop. Adding to that the different distros and fragmenting Hadoop ecosystem, constantly changing versions, etc. - I don't know about you but it can be a nightmare even for analytics (in some cases).
When properly supported by knowledgeable staff with a deep grasp for what it can do as a platform, Hadoop can certainly be and do a lot of things for a lot of use cases.
Hacker News is ridiculously susceptible to propaganda (MongoDB is WEBSCALE! JS framework of the month is the bee's knees!). Please think for yourself and don't read transparent propaganda that is years old. 3 years is an eternity in the big data world.
Though there are very big kinks to work out in getting Hadoop production ready.
So the thesis of the parent article is valid too and not just propaganda.
YARN containers are pretty generic and I have used them to run all sorts of things: Kafka, ElasticSearch etc.
I wonder if it's a typo…