AudioLM: A Language Modeling Approach to Audio Generation
google-research.github.io
google-research.github.io
The paper is here: https://arxiv.org/pdf/2209.03143.pdf
When I read a GPT-3 generated text, I catch myself that I do the same thing. I'm very forgiving. Interestingly, with these speech continuation the illusion of somebody being there on the other side, is much stronger. It's like listening to an old Feynman lecture or the like. You don't quite get it, but surely it must make sense, right. It doesn't of course. The box is empty. Or is it?