I, the writer, am trying to get more ruby people using Pig and hadoop. Feedback appreciated.
This is very useful to me. I've been at LinkedIn for a few weeks and just yesterday I had some time to play with pig and our Hadoop cluster. Ruby is my preferred language for quick hacks, so I'll follow this series.
Good stuff here. As a ruby person, you probably also know about Wukong (https://github.com/mrflip/wukong ) which lets you do Hadoop stuff without looking at Pig. (I don't know about Pig version 0.9.1 but I have had pretty weird bugs on version 0.8)
It looks neat, but it does look like you have to think in mapreduce, which is a big barrier of entry for most people.
Ruby person here. I keep wanting an excuse to try Pig. Is there any reason to do so if my entire data set fits comfortably in memory?
There is if you want to try dataflow programming, and hope your dataset grows to exceed RAM :)
Does anyone have experience doing this in Python stack?
I do, and it is basically the same. Substitute the python avro library for the ruby one, the python voldemort library, and bottle.py for sinatra and you're there.
I am working on implementing this example in Python. There is a problem, however - python-snappy doesn't currently build on OS X. So I can't build the Python Avro bindings :(