1. First and foremost, gorilla arm.[1] My presumption with the "interface of the future" is that it's needed for prolonged use. So, first thing's first, the interface can't be one where our arms require our hands to be higher than our elbows. Unless of course our species got a whole lot stronger in the forearm to support such a feature. Don't see our species doing that anytime soon.
2. Feedback - Right now the feedback loop is eye->brain->hand->brain->eye (repeat) where the hand's pressure against a solid surface is the most important feedback response. With the minority report style interface we currently have a massive delay (comparatively speaking) between the brain->hand->brain loop. We also have to iterate the whole loop much more because we need to constantly assess with our eye where our hand is in 3D (not digital) space. Now let's say the technology gets much better and reduces this to 5ms. We are now bound by the differences of our synapses firing between touch and light. I could be wrong, but it's my assumption that due to the speed of light being the way that it is, that "touch" will always beat "sight" in performance.
For prolonged used applications my bet is on adaptive surfaces. For short term (turning an stove on, flicking a light switch, etc) interfaces I potentially see this Minority Report style interface happening. But does the benefit cost of innovation? Personally I think we are fooling ourselves.
[0] - http://www.ted.com/talks/john_underkoffler_drive_3d_data_wit...
[1] - http://en.wikipedia.org/wiki/Touchscreen#.22Gorilla_arm.22