This should change in next generations of VR helmets.
Purely optical tracking just doesn't seem do it for now, since there are all types of occlusions happening. Maybe something like LEAP with multiple sensors in the room that are able to reconstruct the whole skeletal model up to digits and facial expression (minus eyes, which obviously have to be captured in headset if needed). Currently that is possible with the perception neuron, which is not really fit for casual use based on price (+ USD 1500) and setup time.
I expect to see such high-def video-wall setups achieve market penetration faster than 3D headsets. It's "worse is better" all over again.