They're making a practical distinction that you generally don't have access to the actual thing in an empirical format for which compression will achieve true learning. Instead you have access to training data which represents, let's say, a projection of the actual thing in a smaller space with fewer dimensions.
Like trying to learn from images instead of the 3d world. Humans learn to distinguish between objects in a 3-dimensional space using sight and interaction. This learning generalizably transfers to recognition in 2 dimensions. We don't generally equip models with robotic interfaces to train in 3d before benchmarking them on ImageNet.