Can you provide more details about how you're running them model in-browser?
Then I'm connecting to the device's webcam with navigator.mediaDevices.getUserMedia, and twice a second send a reference to the video stream object to the detector, which runs inferences locally and returns the coordinates of any detected objects.
There's some glue code to do the whole dance to connect to the camera and model, but after that it works pretty seamlessly, on my desktop and mobile browsers