If so, the very action of feeding forward through the layers are hidden reasoning. There is nothing about looping the processing though the same layers a set amount of times, that is any different from copy/pasting the layers and processing it though the same weight. Except it would be stupid waste.
I really don’t understand how this is misunderstood by people that should know better.
Another way to point out the silliness. Raschka's own argument: his Luna vs Sol point shows that ordinary added depth already shifts computation into latents, and nobody called that hiding.