Your criticism is unreasonable. The "cat face" 2012 paper is an excellent paper, and is a breakthrough.
1. They demonstrated a way to detect high level features with unsupervised learning, for the first time. That was the main stated goal of the paper, and they achieved it magnificently.
2. They devised a new type of an autoencoder, which achieved significantly higher accuracy than other methods.
3. They improved the state of the art for the 22k ImageNet classification by 70% (compare to 15% improvement for 1k ImageNet in the Krizhevsky's paper).
4. They managed to scale their model 100 times compared to the largest model of the time - not a trivial task.
You say "it can't be reproduced" and then "can be reproduced", in the same paragraph! :-)
Regarding initializing an input image close to a cat to "get a cat", I think you missed the point of that step - it was just an additional way to verify that the neuron is really detecting a cat. That step was completely optional. The main way to verify their achievement was the histogram showing how the neuron reacts to images with cats in them, and how it reacts to all other images. That histogram is the heart of the paper, not the artificially constructed image of a cat face.