Speaking honestly, the idea is at the level of early experiments. I can get the basic stereo-imaging and motor inputs/outputs within a month. Next would be work on getting feedback system from the environment/surrounding objects including the spectator interactions. This would take another month or a bit more. Finally another month (and more) would consist of experiments on embedding neural nets and scaling the training process. If to move quickly 3 months is enough to build a basis.