We've had "learning-to-learn" algorithms for a few years now. Including LSTMs that can be learn the gradients for other LSTMs, and deep-RL algorithms that can optimize neural networks.
Its hard to say its inventing new things... there is still a clear goal - a loss - and we are optimizing it; poorly in the case of genetic algorithms.
I would say genetic methods are simply poorly described reinforcement learning problems, which means that there is a 1-to-1 mapping between Deep Learning and Genetic Algorithms.
PS. Making a distinction between "Genetic Algorithms" and "Genetic Programming" is like calling Deep Learning "Differential Programming" -- changing the name of a thing does not change the thing itself.
Also note that in the 90s things produced by genetic programming went patented, because they were novel algorithms.
An episode in this case is an iterative call to the Deep RL system until it outputs <STOP> at which step you give it a reward (negative number of collisions say), and before that you can send the actions to your stack machine.
Once you're satisfied with the final number, just concatenate all the produced actions to get your final program.
I just don't see a fundamental difference, especially once you start doing things like Asynchronous Actor-Critic et al. And then you can start doing Monte Carlo Tree Search on top of that since you have a simulator available...