Reconsider whether you really need to: https://www.chrisstucchio.com/blog/2013/hadoop_hatred.html
(although it might look good on your CV)
(although it might look good on your CV)
"Not using a hadoop is like cutting down a forest with a single chainsaw. Using it is like cutting a forest down with unlimited hatchets. It will usually be faster and cheaper with the chainsaw."
I read the link and agree with most of it. All except for when he gets to the 5TB part, but I get that it was written in 2013 and probably not to a fortune 500 audience. If Teradata Cloud could get a reasonably priced entry level tier going, I think they'd be talked about a bit more.
And OP, if you see this, using raw M/R to do things like graph processing is going to be a HUGE pain in comparison to all the libraries built around it.