As someone working on singing synthesis, I know how hard it is to get that last 10% quality that makes a human listener instantly recognise if the voice is real or generated.
These are really impressive results! For anyone interested, my team’s singing work: https://youtu.be/LPy20zSWhZA)