(merged.)
0 karma · joined February 16, 2022
(merged.)
Also, you can write a very small Python script that uses subprocess to call your script. That way you can still use mrjob to set up all your dependencies and handle AWS for you. If anyone is actually interested in that sort of thing, I'd be willing to write a tutorial.
First of all, mrjob has a lot of documentation. Some things are too hard to find in the haystack, but that's more a problem of organization than content. (Always looking for suggestions on how to improve that!) As for the rest...
mrjob tries to make all parts of Hadoop and Amazon Elastic MapReduce jobs easy to write, run, and test. It can run without Hadoop at all, it can use your own Hadoop cluster, or it will automatically handle all the pain of shipping a job to Amazon. mrjob gives you a single API to use across all those platforms and a simple way to switch between them via the command line. It helps you manage file dependencies and greps the myriad logs for errors.
Dumbo (and AFAICT Pydoop) handle the Python-Hadoop bridge, but don't do all that file management and config mishmash. They also don't provide a "Hadoop simulator" for your automated tests. If you want to use one of those tools, you need a Hadoop cluster to run your code, and they won't help you move dependencies and input files around.
The flip side to that is that mrjob doesn't give you the same level of access to Hadoop APIs that Dumbo and Pydoop do. It's simplified a great deal. But that hasn't stopped several companies, including Yelp, from using it for day-to-day heavy lifting. And in my opinion, the simplification is a good thing for most cases.
I should also mention that Dumbo can be faster if you use typedbytes. In theory mrjob could support typedbytes too, but we haven't gotten a patch we were happy with yet.
In the end, mrjob's goal is for these to all give you the same result:
> python my_job.py -r local input.txt > output
> python my_job.py -r hadoop input.txt > output # assuming Hadoop configs
> python my_job.py -r emr input.txt > output # assuming AWS configsI doubt your contacts will migrate unless you do it on the client side.
You succeeded in picking out a minor technical inaccuracy, but not in addressing the point the author was making about rendering performance.
object[0].x = 1
print object[1].x
> 1
Edit: On second read, it looks like you're asking something other than what I thought you were asking. Yes, you could create a list of 100 items and then replace its elements, but that's not idiomatic.Flash was brought up because it was a huge delivery mechanism for a huge portion of casual games over a significant time period. The general theme of the article is that because HTML5 is the up-and-coming technology in the casual games area, many games miss out on basic strategies employed by games on more less cutting-edge platforms.
One real issue with mrjob is that it assumes you're only going to have one key and one value. It isn't straightforward to use multiple key fields. The workaround is to write a custom protocol (which, btw, is very simple [1]) that uses the line up to the first tab as the key, and the rest of the line as the value, probably splitting it on tab as well and passing it through as a tuple. If we had made multipart keys simpler to use, maybe you would have chosen to use a more efficient format.
Anyway, the main part I take issue with is:
"mrjob seems highly active, easy-to-use, and mature...but it appears to perform the slowest."
That's just not true. It would be fair to say that optimizing jobs with multipart keys isn't straightforward and therefore encourages non-optimal code, but that's moot if you're just using one key and one value, as most people do.
I'm really not trying to dump on you here. I liked the post! I would just prefer that it was more precise about these things.
[1] http://mrjob.readthedocs.org/en/latest/guides/writing-mrjobs...
EDIT: If anyone's thinking about downvoting this guy (someone did), don't. This is a discussion in good faith.
Please either mention this difference in your post or update the code and conclusions. If there's a place in the documentation where we should mention optimizations or details like this, I'd be interested to know.
I should have thought of this before. Oh well.
It's just splitting on tab for input and re-joining on tab for output. Some extra logic for lines without tabs.
EDIT: The example in the post is using JSON for communication between intermediate steps, while the Hadoop Streaming example is using a custom delimiter format. So this isn't really a fair comparison; the mrjob example could just as easily use the same efficient intermediate format.
Also, if you're just starting out or want to look over the docs, I'd recommend using the dev version hosted on readthedocs instead of the PyPI version, as the author did: http://mrjob.readthedocs.org/en/latest/index.html
First of all, all of these frameworks use Hadoop Streaming. As mentioned in the mrjob 0.4-dev docs [2] (pardon the run-on sentence):
"Although Hadoop is primarly designed to work with JVM code, it supports other languages via Hadoop Streaming, a special jar which calls an arbitrary program as a subprocess, passing input via stdin and gathering results via stdout."
mrjob's role is to give you a structured way to write Hadoop Streaming jobs in Python (or, recently, any language). When your task runs, it's taking input and output the same way as your raw Python example does, except mrjob is passing it through to the methods you've defined after running each line through a deserialization function. It picks the code to run based on command line arguments such as --mapper, --reducer, etc. The output of the code is again serialized. The methods of [de]serialization are defined declaritively in your class, as you showed.
So why did you find mrjob to be slower than bare Hadoop Streaming? I don't know! In theory, you're running approximately the same code between the mrjob and the bare Python versions of your script. If anyone has time to dig into this and find out where that time is being spent, I would be grateful. Results should be sent to the issue tracker on the Github page [3].
Feel free to ask clarifying questions. I realize I may not be explaining this effectively to people unfamiliar with the ins and outs of Python MapReduce frameworks.
I'm thinking of organizing the 2nd "mrjob hackathon" in the near future, so please ping me if you're interested in contributing to an easy-to-handle lots-of-low-hanging-fruit OSS project. (Particularly if you're a Rubyist, because we have an experimental way to use Ruby with it.)
[1] mrjob is maintained by Yelp, where I worked until recently. It's still under active development, though it's slowed somewhat since Dave and I moved on.
[2] http://mrjob.readthedocs.org/en/latest/guides/concepts.html#...
http://www.winstonchurchill.org/learn/speeches/quotations/qu...
Today's activity has given us a few ideas about how to make those "flyovers" more frequent.
Also doesn't work very well on iOS.
I'm just speculating, but it seems more reasonable than the "because they can charge you for it!!!" hand-waving going on in my sibling comments.