The benchmarks page seems interesting and something I can use to help make an informed decision. Can you talk about how you're measuring some of these? I imagine it needs to involve some human input.
You are right about human input for naturalness, that one we did not automate away with yet. We run blind A/B listening rounds with native speakers.