I think this argument is neither here nor there. For computer vision problems, we use convnets, which are models inspired by a biological model of vision. By doing that we are embedding our preconceived notions of what vision is into our models instead of throwing compute and data at the problem. Earlier attempts using multi-layer perceptrons have been massive failures. Is this consistent with Sutton's analysis or contrary to it?