This feature has been in iOS for a while now, just really slow and without some of the new vision aspects. This seems like a version 2 for me.
This new feature feeds your voice directly into the GPT and audio out of it. It’s amazing because now ChatGPT can truly communicate with you via audio instead of talking through transcripts.
New models should be able to understand and use tone, volume, and subtle cues when communicating.
I suppose to an end user it is just “version 2” but progress will become more apparent as the natural conversation abilities evolve.
If they had released a search engine, which had been suggested, that would be a new product.