This is much more concise than my usual attempts to explain why LLMs don’t “know” things. I’ll be stealing it. Maybe with a different example corpus, lol.
does not answer the general good abstract question and "how semantics possible thru relative context / relations to other terms only?", but speaks to how different modalities of information (e.g. visual data vs. sound data) are likewise represented, modelled, processed, etc. using different neural structures which presumably encode different aspects of information (e.g. layman obvious guess - temporality / relation-across-time-axis much more important for sound data).
tl;dr what you think of as "grounding" is just yet more relative context...