Towards Fully Automated Manga Translation
arxiv.org
arxiv.org
Edit: The other use is to help create better context aware data scrapers that can combine bi-modal information streams and add some body language understanding. I guess it will probably end up in automated surveillance tech like security cameras/mics etc if it works.
It's one throwaway line, but if the line is in Japanese, an American playing an imported ROM might spend an hour of frustration wondering why his sword does zero damage to the final boss.
Actually, i'd say most text heavy games translated from japanese on the nes suffered from this problem and made those games way more confusing than they should have been.
To boot, Simon's Quest also had the atrociously bad idea that NPCs in the game can lie to you and give intentionally incorrect information, making the translation effort that much more confusing.
Then from there maybe I can trigger Cunningham's law and get the attention of someone who knows what they are doing. Sounds like a win-win for me.
UI strings being short usually means hidden heavy context lies in visual elements, so it’ll just strengthen hilarity in mistakes like “Name: SQL Server, Province/Prefecture: Running” (because you know, equivalents to provinces in a region are called “State” in American English...).
“Province” is more or less harmless, but “(has/is/is in/to/like to)Start(ed/ing) type of errors due to missing context can make UI unusable. Oh and it’s un-spottable by non-speakers because they make sense when translated back to original languages.
I didn't find it enjoyable though
Sure, yeah - I think we're on the same page. My post was a reference to enjoyability more than functionality.I have definitely read novels that were long, yet rote and simplistic even in their native language. =)
But they were not works worth reading in any language IMO. They could probably be satisfactorily machine translated (with some human editing) but the result would not be enjoyable except for ultra diehards of the genre who are simply happy to be reading a work from that particular genre, quality of prose be damned.
Those enjoyable xianxia/xuanhuan you mention were either rote and boring in the first place, or they were wonderfully written and had the life crushed out of them by a machine translation that dispensed with all nuance.
If the writing is bad, there's nothing that can really "improve" it other than the original author cleaning it up with the help of an editor. A bad translation is effectively a new work, at best inspired by the original bad script. You could replace a bad Japanese script with a "good" English one - this has happened before - but at that point it's questionable whether any translation has happened at all, you're mostly writing new content inspired by the original work or adhering to broad constraints. What I'd say you're doing here is improving the experience of playing the game, but you haven't done anything meaningful to the writing or script.
In a few cases western companies have licensed Japanese works and spliced them together with entirely new plots for overseas audiences - Robotech is one infamous example where arguably there was nothing wrong with the source material and the result wasn't just a liberal translation.
What distinction are you drawing between working with an editor vs working with a translator? Often it's a very similar process, and there are cases where something is cleaned up in a translation and then that gets incorporated back into the next edition in the original language.
If you can get the output to be 90-95% correct, you can then display the raw and the machine output side-by-side, and have a human make corrections inline. Instead of a team of four working around the clock for a day or two, maybe you could have a translation as fast as it takes to proofread it three or four times end-to-end.
Rev is the same idea in the speech transcription space -- they have humans listen to the audio and fix up a machine-generated transcription.
You can definitely get some of the gist in there, but some of the automated translations are just way off. And pretty much none of them result in good prose. In addition, none of the translations get the names fully correct, so you definitely need someone to go and fix it.
People think of language translation as some sort of same dimension transformation but it’s more like re-projection that involve rotation in upper dimension. Simple warping goes only so far, neural networks give some uncanny slurries, human artists add a lot of their own brush strokes and it’s a lossy process both ways.
Speech transcription is much more straightforward because speakers are supposed to have corresponding single literal expressions for each segments of voice.
In practice if you look at the fan translation community for manga machine-assisted translation is not given much more respect than machine translation - they both produce bad results and in many cases the people who normally welcome even a clumsy translation will reject machine TL and attempt to have it removed, because it often causes people to fundamentally misunderstand the work. The worst cases of machine-assisted TL in manga become infamous to the point of becoming shared memes - try googling "abaj" or "duwang" sometime.
For a very simple pervasive example: Japanese frequently uses gender-neutral pronouns and when translating to English you'll need to appropriately select the right gendered pronoun (or proper name) for each one, if you can. This is something a human can do pretty accurately if they have enough context and knowledge of the material, but it is nearly impossible for a computer to do it accurately without a ton of assistance. In a novel this would be an easier problem because all the necessary context is in the text instead of in the art and panel layouts. You'll note this arxiv paper intentionally cheats on the gender problem.
I think those community translators will be very happy to have some of their work automated.
It's always interesting to see research perspectives on this. I think there's considerable room for improvement, especially given that current translation tech doesn't take into account things like panel layout, horizontal/vertical layout of text, different typeface use, hand-written vs. typeset text, and mixed katakana/hiragana/kanji usage - all things that an author may use to convey tone, implicitly identify the speaker, or add subtext. There are also things with no equivalent in English literature/comics whatsoever, like using furigana to attach a second reading (or dual meaning) to a word.
I can imagine some of this info eventually being pulled in by built-for-purpose MTL tools if a research group puts in the time and energy to do it - the paper appears to be a couple brief steps in that direction because they're feeding in basic information about panel layout and the genders of the people in the panels, but little else. Excited to see what happens in this space even if the idea of even more people trying to translate dense comic prose with Google Translate fills me with dread.
EDIT: I should have done more research before posting this... digging into the authors of this paper, two of them are working at a company that is trying to sell this current technology to authors right now despite the fact that it is not adequate for the challenges of translating comics to other languages. How depressing.
I also can't imagine how it would deal with the idiosyncratic, occasional-nonsensical English that some manga-ka's like to sneak into their work:
http://2.bp.blogspot.com/-Leelb4eEGz0/UzALTIDcgvI/AAAAAAAAEX...
Plenty of paid translation jobs nowadays explicitly hire people to clean up machine translation rather than translate from scratch.
Sure. My point is that it's far from being a taboo practice in the professional translation world. It's a normal practice that is part of a lot of professional workflows.
I'd be super curious if you think this could be a useful tool for translators. Here's a demo video if you want to see what I'm talking about:
2 - 4 Translation
1 - 2 Editing
1 - 4 Cleaning
1 - 2 Typesetting
1 Quality Control
The effort involved in cleaning entirely depends on how good the source material is - you'd think professionals would get pristine high resolution pages without text on them, but you'd be wrong. A high quality cleaning job often means repainting a bunch of the art from scratch. If the cleaning is done poorly the typesetter is screwed.
Some stuff is just plain easy to translate, other works will have a single page full of things you have to look up and sentences that only make sense with previous chapters as context (or even worse, FUTURE chapters as context). In some cases translators also have to first transcribe the work because they're sent blurry jpegs or png files instead of being sent text (again, you'd think this wouldn't happen, but...)
QC and editing can often be done by the same person. One subtle gotcha is that your editor or QC (ideally both, but at least one) REALLY need to know the source language even if they're deferring to the translator - if only one of you is able to check the A->B it becomes very hard to spot subtle problems that can mess up the work.
Human + machine assisted tooling is probably going to be the clear winner for a while when it comes to doing quality translations. As far as fully automated goes, I think it depends on language, but for Japanese it's like 80% there.
All according to keikaku!
(I kinda miss the times when common japanese words were not translated. Current translations often look weird (esp honorifics) and since the rise of Crunchyroll and subsequent demise of fansubs, there isn't much choice left for different translations)
Source https://youtu.be/5rOHpkpYMIM
Or when the translators or fans were weebs. E.g. many Death Note translators didn't translate "Shinigami", even though not only is "Death God" a perfectly valid western cultural concept, the Japanese cultural concept is actually directly derived from the western one.
Two of the biggest names in ripping Crunchyroll have both shut down, one very unexpectedly, and I've noticed a small increase in fansubs since then. I'm kinda hoping for a resurgence now.
https://chrome.google.com/webstore/detail/manga-translator/o...
examples:
- a character who uses clunky or outdated forms of Japanese
- a character meant to have a rural account
- a character with a speech affectation, like adding a cutesy -nyo to the end of their words/sentences
- etc
Perhaps these could be improved if the model was trained on human translations for that specific character and weighted appropriately. With some manga running for dozens or hundreds of volumes that would feasible.
It also definitely wouldn't help with puns, cultural references or other subtleties!
BUT, this could still be really useful as an aid to human translators.
In the worst case even skilled typesetters can spend an entire day cleaning up and typesetting a particularly nasty page if they have to repaint lines and blank out dozens or hundreds of kanji (it happens)
Same with people who ̶t̶r̶a̶n̶s̶l̶a̶t̶e̶s̶ ̶W̶e̶s̶t̶e̶r̶n̶i̶z̶e̶s̶ Americanizes suffixes because I'M PRETTY SURE that EVERYONE calls their classmate MISS or MISTER. And the same goes obviously towards pet/nicknames. Heck, everyone addresses their older sister/brother as big sister/brother and NOTHING ELSE right?
Google Translate translates it more accurately to "rice balls" today - something as simple as that wouldn't be a problem for machine translation.
If it can do visual novels too, I'm in.
to see the English translation over the manga.
:-)