While I think you're right, this is still pretty astonishing.
It's important to note that they're generalizing from a few hours of training data to millions of videos. So the classifier has to be picking up on something deep for it to be re-applied in such a flexible way.
I sort of imagine this approach as being akin to the way Bumblebee (the yellow VW Beetle in the first Transformers) lost his voice, but was able to communicate by switching between radio stations. As that recomposition process being richer and richer, it starts to approximate the real signal...