I love the concept and ease of development, but I can't shake the feeling that the infrastructure is so shaky it almost amount to instant technical debt (sorry if this offends anyone, I'm just a dumb customer.)
I love the concept and ease of development, but I can't shake the feeling that the infrastructure is so shaky it almost amount to instant technical debt (sorry if this offends anyone, I'm just a dumb customer.)
But now Dave is working on mrjob regularly again, hence the pace of recent improvements.
Grandparent is correct about the second-class support for non-EMR production Hadoop usage. Like any open source project, the code only works well if a major stakeholder invests in improving it. Few non-EMR users spend much time contributing, so the situation doesn't improve.
Your performance is going to be complete and utter crap because you're paying for serialization on every single data element.
Dask is higher performance and more pythonic: http://matthewrocklin.com/blog/work/2016/02/22/dask-distribu...
Its free with a permissive license and actively growing.
It is also capable of native HDFS integration, Yarn etc and can do more complex and granular parallel patterns than just map reduce. Also has a API for distributed dataframes and arrays with linear algebra ops.
DISCLAIMER: I don't work for continuum. I just want to see its projects succeed because I was a user will benefit.
However http://discoproject.org/ might be worth a look as a standalone alternative.
Its free with a permissive license.
It is also capable of native HDFS integration, Yarn etc and can do more complex and granular parallel patterns than just map reduce. Also has a API for distributed dataframes and arrays with linear algebra ops.
DISCLAIMER: I don't work for continuum. I just want to see its projects succeed because I was a user will benefit.
But TBH, after a certain scale you should really be asking whether or not you should be using Python.