I can cut vegetables without looking at them. I can use that dynamic to offset the planning and acting phases of my thought process. Falling short of that efficiency will feel limiting.
Maybe X is a button on the keyboard. Maybe X is a gesture.
I can think of some Portal puzzles in particular where timing is important, and you need to hold your aim but wait to click until something happens somewhere else on the screen (so the place you're clicking is not the same as the place you're looking).
I think the same thing applies to e.g. recording network activity in Chrome dev tools. My eyes are on the page to see when the thing I'm interested in finishes loading; my mouse cursor is on the button to stop recording.
It's not a super common pattern, but probably common enough that it would be annoying not to be able to do it.
I am speaking mostly about the desktop interactions. In your Chrome Dev situation, I would look at the cursor before clicking on the stop recording button. I think I might be able to trust the MBP trackpad to do a primed click without looking at the cursor, but I wouldn't trust a traditional desktop mouse to have stayed steady enough.
The mouse is the superior input device. When people who actually need a better input device get one, they get more advanced mice:
I've used a 3D mouse for CAD but am not sure where else it would be helpful?
Nobody wants it for day to day computer interaction. Most people using eye tracking for computer interaction are disabled, because it's a terrible experience.
Oh wait some version of that is built into accessibility on mac already (eye tracking mouse): https://support.apple.com/lv-lv/guide/mac-help/mchl437b47b0/...
Asking me to "blink twice" or anything like that is going to make my eyes lose focus on the screen/content
I'd love to have an eye tracking setup though on a normal desktop computer where I could devote a keyboard key to clicking and I'd never have my hands leave the keyboard.
Maybe they could add other features from just your face though for other platforms, wiggling nose, raise eyebrows, stick out tong, blow a raspberry?
That or they add voice commands, "open", "select" and such.
I suppose ultimately in spatial computing with voice interface and eye tracking the concept of a "click" may die.
We can see this with voice to text, in which despite in theory being so much faster than typing things out, tends to not be due to these details (processing lag, clunkiness of handling different forms such as whether a word is part of a command or should be added to the text).
Same will happen here. People will get over the hype period and then realize hey this isn't actually faster or more efficient than a traditional tool. Apple knows this as well, it's why they're marketing it as a media consumption device first and foremost, where a lot of these problems can be safely ignored.
I had to use it for a while when I was unable to touch a keyboard or mouse while recovering from RSI and I was surprised by how quickly I was able to get to about 80% of my previous productivity using just my voice. I still use it sometimes even though my RSI is fully healed.
But the screens on Vision Pro are larger, and it works great with a keyboard.
So anything that requires larger/more screens seems like it might be more productive than a laptop.