No. You always need some kind of Rosetta stone or other relationship to a known language plus some context and 'plausible guesses' to understand an unknown language. Sure if I gave you
III,IIII;VII — II,II;IIII — VI,II;VIII you would be able to guess that these are elementary number signs in what amounts to a rudimentary table of additions. That much would be true whether the snippet is from a potshard of an ancient civilization or received from outer space via a radio antenna. But outside of context—and nothing would be more out of context than an extraterrestrial culture—you cannot even tell with certainty whether
I stands for 'one' or 'ten' or 'twelve' or 'thousand', and here we've already reached the end of what a text per se can tell you about its meaning if the signs are not clearly pictorial (and even pictorial scripts like early Chinese or Egyptian hieroglyphs are already conventionalized to the degree that for quite a number of signs in either script we are to this day not sure what they depict).
Your idea can not work unless the data that you feed the language model with correlated items. It can't. Imagine I feed a predictor with a long list of images on the one hand and, on the other hand, a long list of randomly ordered image descriptions that may or may not match the images. Do you think you could learn a foreign language that way? You absolutely need the image of a donkey be associated with the name for that animal in the foreign language, and the algorithm is no different.