May not be a popular opinion but I wasn't much impressed by the speech capabilities. It is equivalent to having a ASR + LLM + TTS pipeline or am I missing something?
Yes it is exactly equivalent to that as far as I can see. But it's a big improvement for it to be built into ChatGPT for people who want to use speech input so they don't have to use an external tool.