It's obvious that there are limits to how well machine translation can work unless the models have sensory grounding. I wonder if the problem is that people haven't figured out how to do sensory grounding or that the hardware is still too slow for it to work.