That assumes that the speaker is similar to the person correlating the sounds. For example, if you had statistical data for utterances of English sounds in the context of Magic the Gathering tournaments, and you tried to decipher the speech of a Swahili electrical engineer talking about transistors, you could very well decipher something that's seemingly coherent but entirely incorrect.
It would be an overgeneralization to assume that whales speak about things in the same statistical patterns that humans do.
(I possibly missed a paper)
https://research.google/blog/unsupervised-speech-to-speech-t...
Also, a Youtube doc about researchers attempting to teach dolphins english: https://www.youtube.com/watch?v=UziFw-jQSks
I can’t even understand some other people when they keep switching the target of the pronoun without being explicit.
“He is tired. He dropped the ball on his foot. He yelled at him for being tired.”
(How many people are here?)
compare "dude" in Fig. 1 of https://acephalous.typepad.com/79.3kiesling.pdf
That's fine. The idea is to record them with lot of metadata in situ. Recording what is going on with the whales. (are they feeding? are they traveling? are they in a new location or somewhere they have been for a while? How many wales there are?) And also about their surrounding (sea state, surface weather, position and activity of boats, prey animals etc etc.)
ie train a LLM on English, French, Spanish data. This data only contains parallel text in English-French. Can this LLM still translate to and from Spanish ? Yeah.
There exists no bridge to whale any more than there is aliens from Alpha Centauri.