You’re asking questions about what I assume is the curse of dimensionality as applied to some machine learning problem. The curse of dimensionality is more general: it really boils down to the fact that adding dimensions causes volumes to grow exponentially.
A simple case to visualize that is not related to machine learning is to solve a PDE numerically. Typically this is done via some form of discretization. If you were to solve the PDE in 2D by discretizing a region into, say, a 10x10 grid, then you need to compute the solution at 100 points. If you were to solve the same PDE in a cube, and discretize it similarly, you now need to solve it at 1000 points. Thus, scaling to higher dimensional PDEs rapidly becomes intractable even if the PDE isn’t becoming harder to solve (e.g. resulting linear equations remain equally well conditioned).
In the case of learning, you want your samples to capture behaviors over some region of space. As the number of dimensions increases, so does the volume of space you must gain information from, and thus so does the number of samples you need to explore that space.
The compression you mention is actually a pretty big topic across fields called dimensionality reduction. The idea is that, despite whatever physical process is generating your data (in learning, PDE solving, etc.) giving you a high dimensional vector, this data lives on some lower dimensional manifold (basically you can describe it’s location using a relatively small number of coordinates, e.g. if your data lies on the radius of a unit circle, then a datum is a point in 2D space but you only need to tell me an angle from the origin for each data point). If you can determine this manifold, then any subsequent computation is easier because you can work in a lower dimensional space. The question then is how to find this manifold?
Regarding distances, I don’t quite understand your question, but consider this: there is a difference between distance on a manifold and distance in the Euclidean space in which that manifold lives. For example, the straight-line distance between me and you is not the distance I need to walk to get to you. The former clips through the earth’s crust (thereby leaving the manifold which is the earth’s surface), is a lower bound on how far I need to walk to get to you, and this lower bound becomes a worse approximation the farther apart we are from one another. Therefore, you likely want your algorithm to work with the distance on the manifold as that is more physically meaningful. Finally, distances for mathematical / algorithmic purposes are really relative. You can’t just move the data to be closer together — that would be like setting your height equal to my height in a data base arbitrarily. If you rescale all the distances (convert our heights to kilometers, for example), that doesn’t change anything about the problem once you rescale everything else to respect that change. Finding this manifold and working with relevant distances on the manifold is basically trying to do compression while preserving the most meaningful properties of your data.