Google's image and video recognition? Deep Dream? That's all convolution.
Speech-to-text? That's convolution.
AlphaGo? That's convolution.
Convnets are a great advance in machine learning, don't get me wrong. I hope that soon we get a generalizable way to apply convolution layers to text or music.
There's a few other major techniques (other than convnets) that are important, like RNNs in general.
My concern about the uses of convnets on text that I've seen is that I don't think they can deal with little things like the word "not". (The Stanford movie review thing can definitely handle the word "not", but that's different.) I'm unconvinced as of yet that we're convolving over the right thing. But maybe the right thing is on the way, especially when Google gave us a pretty good parser. And maybe the right thing involves other things like RNNs, sure.
I guess image recognition could have similar cases, and the results just look more impressive to me because I work with text and not with images.
I think there are some papers out of the IBM Watson group on question answering where they use ConvNets. I don't remember looking a the negation case specifically, but Question Answering generally has cases where that is important.
Yet no one solved ImageNet with some convolutions and hand-engineered features. You can't get that performance even if you shift a set of hand-engineered feature across an image in a convolutional fashion. No one solved it with any previous approach. (Are we to believe that CNNs are the very first method to ever try to exploit some spatial structure...?)
So deep networks do bring things to the table beyond hand-engineered features and you are simply wrong.