Why not output MIDI instead, and let the artist manipulate that?
I for one would gladly buy a VST instrument that would generate a constant stream of MIDI ideas, chordings, voicings, variations, from a few seconds of singing or whistling or clapping.
It seems the technology is here to build it, yet (AFAIK) it doesn't exist. Why?
Let's be frank, gen AI projects are ultimately about companies wanting make money by putting small-scale professionals out of work for good. In other words, they want to capture a significant portion of all that capital that currently flows across small companies and projects (even if small scale operations continue a large percentage of their production will require funneling money to Gen AI subscriptions instead of skilled human workers)
This is not true at all, it all started with generating MIDI, and even the very link I shared is of a system that generates MIDI.
I mostly use it as a database of chord progressions, but it can do quite a bit more.