> The synthesis still doesn't know where to place emphasis.True. And yet, these samples aren't in a monotone! That's an enormous improvement.
> You may be unable to distinguish between POOR human voice work and TTS, but not GOOD human reading.
I think these are indistinguishable from AVERAGE human voice work. Keep in mind that the POOR voice work you may have in mind is probably still being done by someone who is, at least nominally, a paid professional.
*> Try e.g. Michelle at: https://cloudpolly.berkine.space/
I'm not able to select a voice other than Oscar (which is definitely worse than this).