Using Euclidean norm over pixels on imagenet will not get you anywhere.
By recasting classification into a kNN lens you're basically (in the large data regime) recasting the problem to kernel learning under some underspecified RBF kernel with an unknown "proximity" metric instead of Euclidean.
In grad school, someone in my lab actually tried scaling kernel methods directly to imagenet too. I don't think that ended up working naively (an interesting neural variant worked for CIFAR10, though [1]).
If all NNs are doing is efficiently learning an appropriate kNN metric (or, more idiomatically and generally put, learning kernel parameters for some implicit neural kernel), that's still really powerful and all we've done is just renamed "learning representations."