There is a big open research problem in the state of the art in visual SLAM (visual SLAM = SLAM from cameras), it doesn't work when the environment is moving (!!!).
Visual slam is still linear-algebra/geometric/keyframe based traditional computer vision (including variants that incorporate GPS/accelerometer info). I think the state of the art is stereo LSD-SLAM, but I could be wrong.