I mean, sight and sound in large language models
are obvious by now, in that rendering them into a token-based representation that an LLM can manipulate and learn from (however well it actually succeeds in picking up on the patterns - nothing about that is guaranteed, but the information is
there theoretically) is currently a conceptually solved problem that will be gradually improved upon.
If reproducing the artifacts and failure modes of human modes of interpretation of this physical data (say, yanny/laurel, or optical illusions, or persistence of vision phenomena) is deemed important, that's another matter. If all that's required is a black-box understanding that is idiosyncratic to LLMs in particular, but where it's functionally good enough to be used as sight and hearing, then I don't see see why it can't be called "solved" for most intents and purposes in six months' time.
I guess it boils down to this: do you want "sight" to mean "machine" sight or "human" sight. The latter is a hard problem, but I'd prefer to let machines be machines. It's less work, and gives us a brand-new cognitive lens to analyse what we observe, a truly alien perspective that might prove useful.