This doesn't change Hadoop from batch processing to real time - it's a new query engine that uses the same data sources/formats as Hadoop and the same interfaces as Hive. So yeah, similar idea to Google Dremel / F1.
It isn't like F1 at all... F1 is a multi-datacenter, transactional datastore and SQL, that google uses to replace mysql.
This aint nothing like that.
You're right that the data is still stored in HDFS, and you can access the data that was originally accessed in Hive. However, because of the way Impala works, it does in fact provide a real-time query interface for data that is sitting in Hadoop (HDFS). The back-end is written in C++ and accesses/processes data much differently than MapReduce which allows for some pretty awesome performance numbers :-)
Correct, the technology is based on Google Dremel.
My reading was that columnar storage was a critical piece of Dremel. "Accessing the same data as Hive" doesn't sound like Impala is doing that. Is that optional then or is this just a query engine?
I can't say for sure, but I believe Doug Cutting's Trevni format is already in the beta but it's a slow version. Impala is initially focusing on architecture, I'm sure columnar format will have attention eventually.