HNHacker News
TopNewBestAskShowJobs

stevejohnson

0 karma · joined February 16, 2022

https://www.hivelance.com/
submissionscomments
stevejohnson··on Mrjob: A Python 2.5+ package that helps you write and run Hadoop Streaming jobs
Good point. https://github.com/Yelp/mrjob/pull/677

(merged.)

stevejohnson··on Mrjob: A Python 2.5+ package that helps you write and run Hadoop Streaming jobs
It's actually pretty close. It's theoretically possible for mrjob to run a script written in any language as long as it supports a simple stdin/stdout protocol. We just don't have any users dedicated enough to implement that stuff.

Also, you can write a very small Python script that uses subprocess to call your script. That way you can still use mrjob to set up all your dependencies and handle AWS for you. If anyone is actually interested in that sort of thing, I'd be willing to write a tutorial.

stevejohnson··on Mrjob: A Python 2.5+ package that helps you write and run Hadoop Streaming jobs
I don't know a thing about Pydoop, but here's what makes mrjob unique with particular respect to Dumbo:

First of all, mrjob has a lot of documentation. Some things are too hard to find in the haystack, but that's more a problem of organization than content. (Always looking for suggestions on how to improve that!) As for the rest...

mrjob tries to make all parts of Hadoop and Amazon Elastic MapReduce jobs easy to write, run, and test. It can run without Hadoop at all, it can use your own Hadoop cluster, or it will automatically handle all the pain of shipping a job to Amazon. mrjob gives you a single API to use across all those platforms and a simple way to switch between them via the command line. It helps you manage file dependencies and greps the myriad logs for errors.

Dumbo (and AFAICT Pydoop) handle the Python-Hadoop bridge, but don't do all that file management and config mishmash. They also don't provide a "Hadoop simulator" for your automated tests. If you want to use one of those tools, you need a Hadoop cluster to run your code, and they won't help you move dependencies and input files around.

The flip side to that is that mrjob doesn't give you the same level of access to Hadoop APIs that Dumbo and Pydoop do. It's simplified a great deal. But that hasn't stopped several companies, including Yelp, from using it for day-to-day heavy lifting. And in my opinion, the simplification is a good thing for most cases.

I should also mention that Dumbo can be faster if you use typedbytes. In theory mrjob could support typedbytes too, but we haven't gotten a patch we were happy with yet.

In the end, mrjob's goal is for these to all give you the same result:

  > python my_job.py -r local input.txt > output
  > python my_job.py -r hadoop input.txt > output  # assuming Hadoop configs
  > python my_job.py -r emr input.txt > output  # assuming AWS configs
stevejohnson··on Mrjob: A Python 2.5+ package that helps you write and run Hadoop Streaming jobs
Oh hey, I work on this. It's come up on HN before, sometimes with misinformation attached. I'm happy to answer questions.
stevejohnson··on Desert Bus: The Worst Video Game
That would be what much of the article is about, yes...
stevejohnson··on DuckDuckGo Changes Homepage Logo, Links to StopWatching.Us Call Campaign
I've noticed that people seem to think DDG is the only good privacy-conscious search engine out there. Lately I've been using StartPage[1], which simply passes through anonymized Google search results. So you don't have the same quality degradation as using DDG, and therefore never need to degrade your privacy for better results with !g.

[1] https://startpage.com/eng/protect-privacy.html

stevejohnson··on Edward Snowden, The N.S.A. Leaker, Comes Forward
Not misquoted, it's directly from the video of the interview.
stevejohnson··on WWDC 2013 Expectations
The only thing you'll accomplish is to get yourself banned.
stevejohnson··on Leaving Google’s silo: Alternatives to Gmail, Talk, Calendar, and more
I'm using hosted.im with my own domain. It works fine, except you can't add GTalk contacts right now, since Google is blocking incoming contact requests from other servers.

I doubt your contacts will migrate unless you do it on the client side.

stevejohnson··on Poll: Do you use your real identity on HN?
There are too many Steve Johnsons in the tech field, but I don't have a better pseudonym, so I try to make my real name count.
stevejohnson··on Why I Develop For The Mac
Technical reasons aside, the canvas tag is still orders of magnitude slower than writing Core Graphics by hand. Personally I'd say almost unacceptably slow, but YMMV.

You succeeded in picking out a minor technical inaccuracy, but not in addressing the point the author was making about rendering performance.

stevejohnson··on Show HN: An HTML5 game we made in 72 hours for Ludum Dare
Interesting and very atmospheric, but I lost interest at about 32 days left. None of my actions seemed to have much effect other than to keep mostly-invisible stats up.
stevejohnson··on Show HN: Firepad, an open source collaborative text editor
I am unable to scroll down on the page without it jumping back up to the animated demo thingy. Chrome 25 OSX.
stevejohnson··on Why Python, Ruby, and Javascript are Slow
That will create a list of 100 instances of the same object.

  object[0].x = 1
  print object[1].x
  > 1
Edit: On second read, it looks like you're asking something other than what I thought you were asking. Yes, you could create a list of 100 items and then replace its elements, but that's not idiomatic.
stevejohnson··on HTML5 Games: Learning from Mobile and Flash Games
You're correct. I intentionally simplified the tenses to avoid having to investigate Flash's current position and potentially misstate it.
stevejohnson··on HTML5 Games: Learning from Mobile and Flash Games
This article has almost nothing to do with Flash, despite the title. It's much more about casual games in general. Using ai2Canvas won't help you implement "gamey" features any more than drinking locally sourced coffee will.

Flash was brought up because it was a huge delivery mechanism for a huge portion of casual games over a significant time period. The general theme of the article is that because HTML5 is the up-and-coming technology in the casual games area, many games miss out on basic strategies employed by games on more less cutting-edge platforms.

stevejohnson··on A Guide to Python Frameworks for Hadoop
I can't find the part of the post where you explain that JSON parsing is the reason for the slowdown. You just say mrjob itself is slower. While I agree that mrjob's defaults encourage the use of JSON, I think it's unfair to blame lack of optimization on the framework, given that the bare Python example could just as easily have used JSON.

One real issue with mrjob is that it assumes you're only going to have one key and one value. It isn't straightforward to use multiple key fields. The workaround is to write a custom protocol (which, btw, is very simple [1]) that uses the line up to the first tab as the key, and the rest of the line as the value, probably splitting it on tab as well and passing it through as a tuple. If we had made multipart keys simpler to use, maybe you would have chosen to use a more efficient format.

Anyway, the main part I take issue with is:

"mrjob seems highly active, easy-to-use, and mature...but it appears to perform the slowest."

That's just not true. It would be fair to say that optimizing jobs with multipart keys isn't straightforward and therefore encourages non-optimal code, but that's moot if you're just using one key and one value, as most people do.

I'm really not trying to dump on you here. I liked the post! I would just prefer that it was more precise about these things.

[1] http://mrjob.readthedocs.org/en/latest/guides/writing-mrjobs...

EDIT: If anyone's thinking about downvoting this guy (someone did), don't. This is a discussion in good faith.

stevejohnson··on A Guide to Python Frameworks for Hadoop
Yep - if it's a Python library that can be 'python setup.py install'ed, then it should be on PyPI. That's true of modules requiring C extensions.
stevejohnson··on A Guide to Python Frameworks for Hadoop
So mrjob isn't slower at all! You just chose to use JSON for your intermediate steps in the mrjob example, instead of the delimiter format you used in the raw Python example. I believe the mrjob docs do say what the defaults are. RawProtocol can be used for intermediate steps just fine.

Please either mention this difference in your post or update the code and conclusions. If there's a place in the documentation where we should mention optimizations or details like this, I'd be interested to know.

I should have thought of this before. Oh well.

stevejohnson··on A Guide to Python Frameworks for Hadoop
The main thing you need to do is publish your Python package to PyPI. This should fill in the gaps in your Python packaging knowledge: http://guide.python-distribute.org/
stevejohnson··on A Guide to Python Frameworks for Hadoop
Here's the code that's doing the actual line parsing: https://github.com/Yelp/mrjob/blob/master/mrjob/protocol.py#...

It's just splitting on tab for input and re-joining on tab for output. Some extra logic for lines without tabs.

EDIT: The example in the post is using JSON for communication between intermediate steps, while the Hadoop Streaming example is using a custom delimiter format. So this isn't really a fair comparison; the mrjob example could just as easily use the same efficient intermediate format.

stevejohnson··on A Guide to Python Frameworks for Hadoop
One last thing I forgot to mention: there was a typedbytes support pull request for mrjob at one point, but it was behind the dev branch enough that it wasn't practical to merge. It is possible (probable?) that typedbytes support will make it into a future release if an interested party could put in the time. I would have done it myself if I had been able to make head or tail of the typedbytes documentation.

Also, if you're just starting out or want to look over the docs, I'd recommend using the dev version hosted on readthedocs instead of the PyPI version, as the author did: http://mrjob.readthedocs.org/en/latest/index.html

stevejohnson··on A Guide to Python Frameworks for Hadoop
Former[1] primary mrjob maintainer here, thanks for the shout-out! I'd like to make a couple of notes and corrections. In particular, I thank you for recommending mrjob for EMR usage, as it's something we've made a point of trying to be the best at.

First of all, all of these frameworks use Hadoop Streaming. As mentioned in the mrjob 0.4-dev docs [2] (pardon the run-on sentence):

"Although Hadoop is primarly designed to work with JVM code, it supports other languages via Hadoop Streaming, a special jar which calls an arbitrary program as a subprocess, passing input via stdin and gathering results via stdout."

mrjob's role is to give you a structured way to write Hadoop Streaming jobs in Python (or, recently, any language). When your task runs, it's taking input and output the same way as your raw Python example does, except mrjob is passing it through to the methods you've defined after running each line through a deserialization function. It picks the code to run based on command line arguments such as --mapper, --reducer, etc. The output of the code is again serialized. The methods of [de]serialization are defined declaritively in your class, as you showed.

So why did you find mrjob to be slower than bare Hadoop Streaming? I don't know! In theory, you're running approximately the same code between the mrjob and the bare Python versions of your script. If anyone has time to dig into this and find out where that time is being spent, I would be grateful. Results should be sent to the issue tracker on the Github page [3].

Feel free to ask clarifying questions. I realize I may not be explaining this effectively to people unfamiliar with the ins and outs of Python MapReduce frameworks.

I'm thinking of organizing the 2nd "mrjob hackathon" in the near future, so please ping me if you're interested in contributing to an easy-to-handle lots-of-low-hanging-fruit OSS project. (Particularly if you're a Rubyist, because we have an experimental way to use Ruby with it.)

[1] mrjob is maintained by Yelp, where I worked until recently. It's still under active development, though it's slowed somewhat since Dave and I moved on.

[2] http://mrjob.readthedocs.org/en/latest/guides/concepts.html#...

[3] http://www.github.com/yelp/mrjob/issues/new

stevejohnson··on The great California Exodus
This quote doesn't add to the discussion and is also incorrectly attributed to Churchill.

http://www.winstonchurchill.org/learn/speeches/quotations/qu...

stevejohnson··on Buildy, a real-time HTML5 building game - come build in our biggest world yet
You are a very good guesser.

Today's activity has given us a few ideas about how to make those "flyovers" more frequent.

stevejohnson··on Buildy, a real-time HTML5 building game - come build in our biggest world yet
If you're in the world and someone is doing this to you, please say something in chat and we'll ban them.
stevejohnson··on Buildy, a real-time HTML5 building game - come build in our biggest world yet
If you want to get updates about Buildy, you can follow our blog:

http://blog.playbuildy.com/

stevejohnson··on Buildy, a real-time HTML5 building game - come build in our biggest world yet
NB: We've only tested on recent Chrome and Firefox builds on OS X. Bug reports in other browsers are appreciated, but I can't promise we'll be able to act on them in the near future.

Also doesn't work very well on iOS.

stevejohnson··on Pattern - Web Mining Python lib
NB: screen-scraping Yelp is against the TOS and you'll get shut down pretty fast if you try it.
stevejohnson··on Apple Officially Reveals The iPhone 5
Likely for size reasons. "Micro" USB is still pretty thick on the iPhone thickness scale.

I'm just speculating, but it seems more reasonable than the "because they can charge you for it!!!" hand-waving going on in my sibling comments.

← PreviousPage 4 of 14Next →