This approach would be useful if 3D devices let you render a view for each eye. And report head and hand movements as events. So one would build a whole 3D engine with the current Browser APIs.
The approach I was thinking about was to tell the 3D device "This object is located at x,y,z position 123,40,8". And the device does the rendering. It's probably much faster, as the device probably has a lot of hardware optimized for 3D rendering. Let alone things like AR, where you have to analyze a given video input, figure out that there is a real table in a certain position in space where you can put 3D objects on, calculate the physical interaction and shadows etc.
Not sure which approach is better. Time will tell.
https://github.com/pmndrs/react-xr
Apple announced support for WebXR on VisionOS as well.
I think it will be a lot more interesting once webGPU hits too, as it will be closer to native-level GPU programming but portable between both native and web contexts.