I once asked a similar question on some online forum [1] where many linguists hung out. My question was if an English-only speaking household left a general interest Spanish language TV station on most of the time when they weren't actively using the TV to watch something, so that their child received a very large exposure to Spanish language programming (news, sports, soap operas, sitcoms, movies, etc) from birth onward, would the child naturally learn Spanish?
I don't recall for sure what the linguists who responded said, but I think they all said the child would not learn Spanish from this.
[1] I have no recollection of where this was.
If the child will actually watch the Spanish TV he will learn the language.
EDIT: Even now I often learn new japanese words (and remember them) just by watching animes. The difference is that now I have english subtitles but back then I had no subtitles, only the images to help me understand the meaning.
Not knowing basically anything about AI state of the art what stops us from feeding a RNN image data and text data and make it correlate them automatically by context? Just like a child learns words by hearing them many times in similar contexts so could a RNN.
I imagine the biggest problem is gathering and structuring the data. We humans receive lots of data and have lots of time to process it in our lives compared. And by lots I mean difference of a few orders of magnitude. It's amazing what this thing learns in just a few hours of processing.
I've picked up quite a bit of Russian by watching Discovery channel this way.
The methods involve providing more detailed feedback at each example. With most training data used now, we give a 0 or 1, does this example belong to this class. In the teacher networks, they were able to teach with more subtly: this is definitely not a car, it is very lizard like and a little snake like.
https://en.wikipedia.org/wiki/Language_acquisition#General_a...
Although for obvious reasons this is very hard to study experimentally:
https://en.wikipedia.org/wiki/Language_deprivation_experimen...
Seems it would be far harder to infer the basic initial structure from just plain text.