https://www.theregister.com/2023/02/18/vr_telemetry_identity...
The example described in the paper in your link is using cameras that are setup to provide very high resolution images of people. It would be like using full-frame portrait images for face recognition, and then expecting that to translate to real-world scenarios where you might only have 20 pixels on a face, and the person is off-axis to the camera.
Gait detection has been discussed for a while, and may definitely be a thing one day, but right now we are barely at the point where pose estimation is a thing in security video. Very far from being able to do high precision pose recognition and sampling over successive frames to model something that would qualify as "gait".