I guess the idea will be somewhat similar, going from coarse to fine details, such as for 3D structures.
Maybe the original author benanne could give his insight.
Maybe the original author benanne could give his insight.
That said, the gap between perceptual modalities (image, video, sound) and language is quite large in this regard, and probably also partially explains why we currently use different modelling paradigms for them.