I think the author's biggest problem is the bottom-up approach of creating lots of raw clips and then at a higher level trying to assemble them. I've found music generation to consistently be a very similar task to text generation, and to me the author's approach seems analogous to generating lots of random sentences and then trying to assemble them into a short story.
Everything I've done that's worked has been completely top down - the core of the engine determines the state of the song (we're in the A section, bar 7, the current chord is V/V, etc.) and the low level parts generate something to match all that. E.g. the melody part might decide to play the root tone of the current chord, but it has to ask the core what note that would be.
The hard thing about this approach is that you need an environment where you can dynamically play notes of an arbitrary pitch/length/instrument, which either means lots of sound fonts or realtime audio gen. I've been doing this with WebAudio, and while it works it's been very painful - for me the "play sounds" part of procedural music has more difficult than the "compose music" part.