The problem is that this is incredibly computation intense, since the number of possible programs is huge. Right now it's viable for improving small code parts with a know-correct starting point. Maybe some day computers become fast enough to make more viable.
The challenge is that 1) doing it for non-trivial functions and getting results faster than much simpler generative methods that depends on heuristics is a really hard problem (but can work better when we don't have reasonable heuristics); 2) writing exhaustive tests for a lot of the problems we care about is likely to take more effort than writing the code in the first place.
I think if you want something like this the effort is best expended on tools that help you create a consistent, concise and exhaustive model / test-suite rather than code to implement it, with a focus on making it possible for a human to read and sanity check the generated model.
In this case, the tool is basically trying to create a model that matches the tests, it's just that it never makes the model explicit other than in the form of finished code, which prevents us from verifying that the model is correct other than by inspecting the code and/or expanding the tests.
For small functions that might be helpful, but for larger pieces of code, I think it is likely that generating the code directly is likely to lead to code that is near impenetrable and impossible to validate expanding the test suite. E.g. a recurring problem of research in genetic programming has been that a lot of the resulting solutions are hard to understand even very small/simple algorithms. And for bigger problems it's not unusual to end up exposing weaknesses of the fitness function rather than solving the intended problem.
It's still an interesting project. I just think we're really far from having something with wider appeal.
That said, webyrd has shown kanren embedded lambda calc (evalo relation) to find which program would be reduced to some value..