>> What seems novel here is how significant architecture alone is (I wouldn't guess it being this powerful), and this possibly has implications on the connection with the human brain -- in which my vague impression is that connections themselves are much more tuned than weights; also possibly giving greater emphasis on architecture training.
You are surprised by how significant architecture is. I am not surprised at all. I think it's been very clear for a while now that the success of deep neural networks is entirely due to their architectures that are fine-tuned to specific tasks. On the one hand, it's obvious that this is the case if you look at the two major successes of deep learning: CNNs for vision and LSTMs (and variants) for sequence learning. Both are extremely precise, extremely intricate architectures with an intense focus one one kind of task, and that kind of task, alone. On the other hand, the vast majority of neural net publications are specifically about new architectures, or, rather, tweaks to existing architectures.
In fact, in my boldest moments I'd go as far as to suggest that weight training with gradient optimisation has kept deep neural nets back. Gradient optimisation gets stuck in local minima, and it will always get stuck in local minima, and therefore always be dumb as bricks. Which is evident in the way deep neural nets overfit and are incapable of extrapolating outside their training datasets.