We've been using the FLUERS eval and you can see comparisons to other models on the market in the post
https://openai.com/index/introducing-our-next-generation-aud...
Curious if there's a benchmark you trust most?
Curious if there's a benchmark you trust most?