If anything, I've become more convinced that text alone has enough information to solve some eerily specific problems. It's quite a powerful tool.
What? Even I couldn't answer that. I thought it was a brainteaser. "Water" in common parlance refers to its liquid form, as opposed to ice or steam.
Theoretically you could get there but you'd need to connect a lot of dots in your algorithm. Throwing more text at existing algorithms won't reliably answer this kind of question.
Aren’t people working on combining language and visual models now? It will be interesting to see how those do.
A good starting point is a system that has a pretrained list of objects it can recognise. For example, we can train our image model what a book is, so when a user asks for a book, we can correlate the "bookness" feature (trained part of the output vector from the vision system) to it 1-1, and connect the bounding box from the vision system to the named entitity "book". However, say a user instead asked for a "novel". While this is also a book, we need a way for the two to be connected. For this, we use a quite different type of language model to connect the named entity "novel" to a list of known objects, one of which is "book". It does this by scoring how similar words are in the language model, and picking the closest match that is in the list of known objects. Over some threshold of similarity (measured by distance in abstract word-vector space), it is determined that this is a new, unknown object. What happens next varies, but is usually some variation of the system guiding the user to teach it what the object is that they meant (something like: point to the object, or put the object on top of a special evenly lit blank background for the visual system to learn it's features), so it is known for next time, and given an identifier from what the user called it.
This is the most basic form, but (as with everything) it gets more complicated. Two areas I am looking into at the moment are past reference language learning, which is about learning new visual features from users in operation. For example: user refers to this particular book as an "open textbook", later asks for a "closed textbook". Model knowns "open" and "closed" are mutually exclusive, so looks for a second match for book that is dissimilar in some visual features to the first, and learns some weak relation in those features to "open" and "closed", possibly to be reinforced later by repeated user expressions.
The other area I am looking at is learned spacial relation grounding. That is, trying to be able to resolve something like "The third book from the left" or "the book on top of the blue table". This is interesting because it breaks a lot of current grounding approaches, in that it requires some recursion (you can chain spatial descriptors in natural language e.g. "The book on the table in the room in the leftmost house on the hill...") and self-referential techniques. Ways of doing it aren't as well researched, and we're working on a method for doing it atm :)
Say someone asks your model to find pictures of a leash. If it’s never been trained on leashes it can use the language model to know they’re associated with dogs, attach to the collar, etc and actually find pictures of leashes and start training itself.
The possibilities are endless. Let me know if there is further reading I can do.
Grounded Language Learning: Where Robotics and NLP Meet*