This is actually not as cut and dry as it sounds. Generally speaking the difference between VO and SLAM is how it handles accumulated errors to do LC/relocalization. Almost by default a robust VO system using some kind of sensor fusion does a form of relocalization but rarely loop closure.
I say this because I feel like it undercuts Apple's effort by stating it's a purely VO approach and I don't want people to get the impression that you can't do round trip SLAM without active depth (IR etc...).
"Almost by default a robust VO system using some kind of sensor fusion does a form of relocalization but rarely loop closure." This is not true, Visual(Inertial) Odometry systems typically estimate "odometry" i.e frame to frame motion tracking hence the name, such as in [3] and [4], whereas SLAM systems, such as [5], store a map of features, and their descriptors to relocalize which isn't necessarily a typical feature of pure VO systems.
[1] https://developer.apple.com/arkit
[2] https://developers.google.com/tango/developer-overview
[3] https://github.com/uzh-rpg/rpg_svo
So the reason I say that it's "almost by default" is because on really quality VO systems, if you turn in a 360 degree circle, the image descriptors and pose estimates returned are almost exactly the same. It's not of course the same, but the effect for an end user is the same and can be extended relatively easily.