Residual neural networks[] employ skip-layer connections. It's thought that this helps prevent gradient decay in deep networks. It's a very important advancement.
Resnet was when we figured out how to throw a fuckton of perceptrons together and make it actually work. We didn't think networks scaled before that.
iirc that was AlexNet, ResNet came 3 years later whose refinements ended up making it better than a human at object detection.
I think the argument is that before ResNets, the depth of the network was constrained and could not be scaled easily. With ResNets (and of course Highway Nets; hi Jürgen!) the depth is just another hyperparam.
In between those, VGG was parameterized by layers, though severely limited in scope (16-19 layers).