Even visual odometry and SLAM, necessary for AR, can be solved through DL.
I, for one, can't wait to have something in a raspberry Pi form and power factor that can run a big deep CNN at 30fps.
That opens the way to a lot of neat gadgets that can be summarized as "computers that can see":
https://www.ben-evans.com/benedictevans/2019/7/19/computers-...
There are a lot of very sophisticated and computationally intensive algorithms out there that can use specific sensors or combinations of these sensors to infer other useful things.
Now take into account you may have multiple phones collecting data distributed in some spatiotemporal fashion, you can really get fancy with some aspects and increase accuracy of some of these algorithms.
Instead of hand waving, one novel example I've worked with is using video feeds from streams or rivers to perform LSPIV (large scale particle velocemetry) analysis to estimate instantaneous discharge rates of rivers. There are at least 6-7 other use cases for projects I've been involved with that could leverage these resources in interesting ways--some use phones, some use other sensors.