Check out Hive - https://cwiki.apache.org/confluence/display/Hive/Home - it uses a "SQL like" syntax to let you run ad-hoc queries across data in a Hadoop+HDFS cluster. We're using it to run reporting nightly on ~10gb of data and are pretty happy with it.