What I really want to know is what kind of robot arm motion is produced when the network is given a cat image to classify. More specifically, what kind of insights has it learned from one control domain that it then applied to another?
I imagine that the simulated 3D environment and the actual control of the robot arm must have some degree of interconnection neurally.