I wonder if the author of the article knew about the series, or do they both just independently came across this topic to talk about it.
The extraction of features from a corpus, the features significant to certain solution, is always and since day zero - compression. As this is the definition of compression - efficient and potentially lossless feature extraction.
So they're both sourcing a bit broader zeitgeist.