Bit late to the action, but do I understand correctly that this trains a harmonic synthesizer? That would explain why the reverb in "How does the amount of training data affect synthesis fidelity?" sounds a bit off, as it would get synthesized on unrelated input.
But what I wanted to ask: is the final model actually understandable or visualizable, e.g. as an envelope per harmonic that depends on F0, loudness and F0 confidence?
Anyway, I'm impressed.