[edit] Some of the quotes from his 2006 presentation are quite relevant to this...
... "Now, multi-touch sensing isn't anything- isn't completely new, I mean, people like Bill Buxton have been playing around with it in the '80s." ...
... "Now this is a photographer's light box application. Again, I can use both of my hands to kind of interact and move photos around. But, what's even cooler-
(uses fingers to 'grab' two corners of one of the photos and 'pulls' it to full screen size)
is that, if I have two fingers, I can actually grab a photo and then stretch it out like that really easily. I can pan, zoom, and rotate it effortlessly.
(slides piles of photos around)
I can do that grossly with both of my hands,
(pulls photo out of stack & enlarges it)
or if I can do it just with two fingers on each of my hands together.
(grabs empty space around photos & zooms in and out of canvas)
If I grab the canvas I can kind of do the same thing- stretch it out- I can do it simultaneously, where I'm holding this down-
(holds pile of photos down while pulling out another)
-and gripping on another one, stretching this out like this.
Again, the interface just disappears here. There's no manual. This is exactly what you kind of expect, especially if you haven't interacted with a computer before." ...
Which sort of begs the question that if an expert in the field thinks the gesture is exactly what you would expect, even if you had no expertise whatsoever, then how does that not qualify as obvious?