The "warbling" factor stood out to me. It's like you could hear the interpolation between audio samples.
Yeah, there was some strange timing between sentences and words at times.
What I am curious about is how much selection was involved in the fake ones. Like is it the cream of the top of the fake ones that sound the most natural or is it a typical output?
I noticed that, and also that the spacing in words was too uniform.