Your big data toolchain is a big security risk
vitavonni.de
vitavonni.de
The other often misunderstood problem with Hadoop/Spark/... is that the security model is basically the same as NFS.
If you don't use Kerberos any user with access to the Hadoop cluster has at least full read-access and likely even write access if he has superuser rights* on the client machine. => http://www.openwall.com/lists/oss-security/2015/04/16/21
If you have important data on your HDFS and you did not put everything behind a thick firewall and you are not using Kerberos you have a problem. You'll likely still have a problem...
* This should also work without superuser rights, Hadoop just takes the username from the client.
The article lays its points clearly, but I'm unsure if "don't use hadoop" is a viable conclusion - and I'm sure the Hadoop community knows about its faults better than anyone.
A "three year old" long term support release. Isn't that what you wanted?
In most cases hadoop is total overkill -- developers don't understand how to reduce problems, might have never heard of sampling, good data-structures, know your problem, etc.
For sure, there are valid use scenarios for map reduce/hadoop, etc. but in many cases it's a big waste of money.
"We don't do big data, we do little data" -- said no dev team ever, since Big Data has appeared.
Loading a 100MB Excel file -- big data. Dumping a few GB to SQLite -- big data. And programmers are now of course "Data Scientists".
Same with services. When microservices became cool, all the other services have disappeared and everyone is doing microservices.
And I would disagree that Hadoop is a big waste of money. There isn't anything else really that comes close to its cost effectiveness once you reach a certain data size.
you should also probably have a model cluster to allow you to experiment with up grades.
Exactly.
there may be more substance later but once I hit the bullshit about "iFanboys" I decided not to bother. At best I'll just shake my head wondering why people think their personal anger is a compelling argument,
However, calling out iFanboys for particular scorn detracts a bit from the article's value.
They might have a point about using packages instead of "curl | sh", but I'm not sure I take them totally seriously.
This might start conversation, at least?
BTW I used MR back in the 80's on the then largest prime computer cluster for BT
Use Sonatype's repo then? http://central.sonatype.org/pages/requirements.html#sign-fil...
I want to be able to build my entire project hierarchy from source.
Sonatype prefer if you provide sources, but they don't require you do so, also there doesn't seem to be any way of actually verifying that the source matches the binaries.