From the perspective of 1960, early "robots" were expected to be physically capable and mentally feeble. Good at logic. Capable of making a sandwich. Weak at empathy and whatnot.
Even from the perspective of 2025... most people don't understand how slowly robotics has advanced. Human-level performance at laundry folding remains a distant dream. Empathy is increasingly trivial.
So Siri.. and voice UIs generally. The bottlenecks have been in unexpected places.
In general, we just don't have very good UI paradigms for voice. Voice recognition is finally good. LLMs theoretically add a lot of capability. But... there just isn't a great UI.
It's like trying to use a smartphone with a nipple mouse instead of touch. You can slowly hack your way to making specific tasks/features work... but there is no radiation event where lots of tasks become possible.