Taste and olfactory are matters of chemical compositions. It will take an incredible effort but something similar to a mass spectrometer can be used to detect every taste and smell we can think of and beyond. How fast and how efficient they can be is probably the main challenge.
Touch is difficult. We don't even know fully why or how does an itch "work". But force, temperature, atmospheric, humidity sensors, etc are widely available. They can provide a crude approximation, imo.
Just off the top of my head. I am sure smarter people can come up with much more suitable ways to "embody" a machine learning model.
If reproducing the artifacts and failure modes of human modes of interpretation of this physical data (say, yanny/laurel, or optical illusions, or persistence of vision phenomena) is deemed important, that's another matter. If all that's required is a black-box understanding that is idiosyncratic to LLMs in particular, but where it's functionally good enough to be used as sight and hearing, then I don't see see why it can't be called "solved" for most intents and purposes in six months' time.
I guess it boils down to this: do you want "sight" to mean "machine" sight or "human" sight. The latter is a hard problem, but I'd prefer to let machines be machines. It's less work, and gives us a brand-new cognitive lens to analyse what we observe, a truly alien perspective that might prove useful.
No matter how you build it, it is still experiencing everything a human can experience. There's just no guarantee it would react the same way to the same stimuli. It would react in its own idiosyncratic way that might both overlap and contrast with a human experience.
A more "human" experience simulator would paradoxically be more and less authentic at the same time - more authentic in showing a human-style reaction, but at the cost of erasure of model's own emergent ones.
Apart from that, I'm afraid that at this point research on sensory input apart from audio and visual needs much more advancement. For example, it's not clear to me what kind of data structure would be a good fit for olfactory or sensory training data
Touch and such can have some approximation done through various sensors like temperature, force, humidity, electromagnetic, etc.
Don't get me wrong I would be curious to see such research done to see whether it would improve anything above the stochastic parrot level - it's just going to take a while to figure out what is even relevant
But an LLM has no problem at all deciphering and processing and most importantly, responding meaningfully to all the ways we can use or encounter the word "lie". I contend that if a model large enough is trained on enough data, the concepts will automatically blend and explain each other sufficiently, or at least enough to cover your example and those similar to it.