This is a bit different. These audio clips use the default voice of each of these systems. I was asking about zero-shot voice cloning, i.e. transferring a recorded voice and synthesizing speech in that voice.
I tried zero-shot voice cloning in all of the top OSS models in the Arena and performance was bad.