It looks like lately a lot of progress have been made in audio generation / audio understanding (everything related to speech, I mean).
Is this related to LLM, or is this a completely different branch of AI, and is it just a coincidence? I am curious.