No, they require training a deep neural network first. Then training the shallow net to mimic the deep one.
The claim is that the shallow net can have the same number of parameters as the deep net. It's always been known that a large enough shallow NN can theoretically approximate any function. But that it can do so with few parameters is very surprising.
My best guess is that the majority of parameters in deep NNs are unused or redundant. I.e they have a low weight, or don't influence the output very much, or another neuron computes mostly the same function. Many types of regularization like dropOut and weight decay heavily encourage this.
Whereas the shallow NN is forced the maximize the usefulness of every single parameter to get the same result. It doesn't have to worry about overfitting. I am not sure if the paper accounted for this though, it's been awhile since I read it.