The "Compression" part is indeed about the model finding order in the training data. But this is not about applying some predefined compression algorithm on it, but by actually learning an algorithmic representation of the data.
This has nothing to do with "consulting the training data", because the training data cannot be reconstructed based on this information.
Roughly speaking, there are two types of information being stored in the NN: Memorization and generalization, or, shannon entropy and kolmogorov complexity.
This paper is highly underrated: https://arxiv.org/pdf/2505.24832