Search Big Data in Five Minutes with Pig, Wonderdog and ElasticSearch
hortonworks.com
hortonworks.com
It's just 450MB of data (in Avro format). Why should the filter be "slooooooooowwww" ?
Compare this to Cascading which is far easier to unit test.
Also PIG is a hog when it comes to making jars. If you don't install pig on EVERY node in the cluster and rather provide it via a submitted uber jar, the jar is HUGE (50 or so megs). PIG's dependencies are ridiculous.
Again compare to Cascading who has a far smaller foot print.
I would like to a strong comparison of HBase. Pig is useful for one-off, ad-hoc stuff, but I don't think it's production ready.
We unit test our pig scripts, it's pretty straightforward given MockStorage class we have contributed (you can find it in pig trunk). Granted, we've long been separating load statements from actual logic, which allows us to fairly easily mock up data to feed into the Pig flows. It would be harder if your loads and flows are in the same file.
https://github.com/linkedin/datafu/tree/master/test/pig/data...
Our tests take seconds, and can run from eclipse (we use Pig's local mode).
You sank 3 days into modifying internals of PigUnit instead of just parametrizing your loader statements?
Testing is a good point. If you're running Pig through their APIs it is definitely easier to test than command-line running scripts. We've written test code that reads and runs pig scripts through the API using fixed sample data (stored in HDFS for easy access), read the results, and compare it to expected results (also stored in HDFS). Honestly, you don't need too much input data to prove the correctness of the query.
Remember that Pig also supports placeholders in your scripts so you that you can set them in run-time to define input/output paths, etc. This makes testing easier.
Dependencies can also be stored in HDFS which makes it simple to run your scripts w/o the need to distribute jars around.