Sure it's nice to have a display you can always see while keeping your hands free, but even when it perfectly recognizes what you say voice input is still slower and clunkier than using your hands. Gaze tracking is similarly clunky compared to multitouch.
Not to mention the privacy aspect (would you want everyone around you to know exactly what you're typing/searching for/interacting with?)
You could have some kind of physical input device you keep in your pocket - but at that point you're carrying around more stuff than if you just had a smart phone. (Edit: sensors on your fingers so you can type on an imaginary keyboard?)
Basically it just doesn't seem to make any sense to me, but hey maybe I'm missing something. Useful tech for certain use cases sure, but I don't think those use cases really overlap enough with what we use smartphones for to replace them.
[1] https://www.reddit.com/r/sciencefiction/comments/2ige0f/im_a...
Or type on some sort of virtual
[2] https://en.wikipedia.org/wiki/Chorded_keyboard (Octima) combined with something like
[3] https://www.openstenoproject.org/plover/
combined with something like
[4] https://en.wikipedia.org/wiki/T9_(predictive_text)
or by 'grunting/subvocalizing' into some tiny pearl fitted to your neck/larynx (from the outside OFC).
Voice is also much more distracting than tapping thumbs, as anyone who works in an open floor plan office can attest.
The real challenge with XR is the vergence-accomodation conflict, it's solved in AR but not VR [0].
AR input software is a commodity, open-source versions exist. They use gesture detection.
But if you consider this to be AR, how is it different to XR and how can it be solved for AR but not for XR?