Someone shoot me down if I'm wrong:
In this model (one of many infinitely many possible models) a word is a point in 4096-space.
This article is trying to tease out the structure of those points, and suggesting we might be able to conclude something about natural language from that structure - like looking at the large-scale structure of the Universe and deriving information about the Big Bang.
Obvious questions: is the large-scale structure conserved across languages?
What happens if we train on random tokens - a corpus of noise? Does structure still emerge?
It might be interesting, it might be an artifact. I'd be curious to know what happens when you only examine complete words.