Music generation is one of the easiest ways to "spook the normies" since most people are completely unaware of the current SOTA. Anyone with a good ear and access to these tools can create a listenable song that sounds like it's been professionally produced. Anyone with a good ear and competence with a DAW and these tools can produce a high quality song. Someone who is already a professional can create incredible results in a fraction of the time it would normally take with zero budget.
One of the main limitations of generative AI at the moment is the interface, Udio's could certainly be improved but I think they have something good here with the extend feature allowing you to steer the creation. Developing the key UI features that allow you to control the inputs to generative models is an area where huge advancements can be made that can dramatically improve the quality of the generated output. We've only just scratched the surface here and even if the technology has reached its current limits, which I strongly believe it hasn't since there are a lot of things that have been shown to work but haven't been productized yet, we could still see steady month over month improvements based on better tooling built around them alone.
Text generation has gone from markov chain babblers to indistinguishable from human written.
Image generation has gone from acid trip uncanny valley to photorealistic.
Audio generation has gone from 1930's AM radio quality to crystal clear.
Video generation is currently in fugue dream state but is rapidly improving.
3D is early stages.
???? is next but I'm guessing it'll be things like CAD STL models, electronic circuits, and other physics based modelling outputs.
The ride's not over yet.