Just a variational autoencoder with a discrete latent space is good enough to learn a usable phoneme recognition and pronunciation model from raw WAV files with unsupervised learning.
And clip shows is just for far you can get without supervision. So what's the point in paying for artificial data if you can solve the problem without it?