Yeah, totally agree that something related to this will likely be the next paradigm. I've been putting together experiments in different directions trying to find that thing that's missing but haven't really found a killer use case yet to pull it all together.
That's a really cool idea that once you can get something somewhat reliably consistent generated, you can kind of let your A/B tests start to run themselves with just rough guidelines on what you're trying to optimize for...