It's definitely feasible and widely used for reinforcement learning on low dimensional systems, where the neural networks are small and the simulator is more expensive than backprop. On other hand, deep Q learning from pixels on Atari is practically impossible without GPUs.