Funky Fantasy IV: A Machine-Translated Video Game Experiment
legendsoflocalization.com
legendsoflocalization.com
I suspect they translated each text box separately, instead of joining multiple text boxes into one string - because I frequently get much better translations than this from deepl. Japanese is a language that is often difficult to understand without context, so the smaller the chunks of texts are the worse a computer will do.
Stuff like: "Your brother is probably a man. Probably an adult." is just a bad translation of でしょう.
"Put the bomb in the ring in his hand" The second part is an overly literal translation of 手に入る, and the ring was probably called 爆弾の指輪 or something
”Very!" is probably just 全く
etc
These are things any intermediate learner would know, and yet AI models still struggle with them.
Eg, with DeepL:
ダークエルフなら ほくとうのしまの どうくつに すんでるよ! => If you're a Dark Elf, you'll find me in the caves of the Far East!
ダークエルフなら北東の島の洞窟に住んでるよ! => If you're a Dark Elf, you live in a cave on the northeast island!
DeepL also seems to be quite sensitive to the placement of spaces.
手に入る literally translates to "put into hand", but it means "acquire" in the context of opening a chest and finding a potion. Japanese games say "potion put into hand" and English translations say "You received a potion".
全く literally means "truly" or "very" but a lot of the time people use to mean "seriously" in a sarcastic sense. Like, "seriously!?" or "jeeze!". A beginner would think "very" but somebody with experience would know they're just complaining in a way that doesn't elicit* a response.
"Brother" seems to be coming into it because Rydia is addressing Edward as "o-nii-chan". Pretty classic example of something that you couldn't translate accurately without knowing who the characters are.
I think the project have been from Ravi Purushotma: http://langwidge.com/theoretical.html
BBC News article: http://news.bbc.co.uk/2/hi/technology/4182023.stm
The neural network result looks better, but is less accurate, yet much more playable? That's an odd conclusion...
Version 1 is word-salad, but the word salad usually includes the subject and object of the sentence.
Version 2 pretty much always gets a feasible sentence, but sometimes the sentence is completely unrelated to the original.
Which one is more playable is a bit subjective. DeepL looks quite a bit better than either of the google-translate versions though.
The DeepL version definitely looks far better than 1 or 2.
https://m.youtube.com/watch?v=WnzlbyTZsQY&feature=youtu.be
Would totally play this mod.
ダークエルフなら ほくとうのしまの
どうくつに すんでるよ!
https://youtu.be/87eL0K9Ld8s?t=79AFAIK the premise of NNMT is that it doesn't try to do much in the way of parsing, but tries to "learn" the association between phrases in the two languages by matching of the training data, so if it sees "fitness room" and "Thailand" and that happened to somehow align with an occurrence of ほくとうのしまのどうくつ , that's what it will think they translate to.
You can see some interesting "probing of Google Translate's guts" here:
The original translates to something along the lines of "Dark elves live in the caves on the islands in the north east".
Maybe it took "north east" to mean Thailand and instead of separating しま and の it just took it as one block of text しまの (Shimano) like the bike company, and therefore fitness and どうくつ is cave which is in a way a type of room but that is one hell of a tortured translation.
Good times.
But GPT-3 still hasn’t seen “much” non-English test (they only intentionally feed it mostly-English-language sources for now) so it’s definitely not as good at translation as a GPT-4 that had seen equal amounts of other-language corpus would be.