Cutlet: A Japanese to Romaji Converter in Python
dampfkraft.com
dampfkraft.com
The annoying thing is that the Japanese are equally quilty of this. I live in Finland so Ä/Ö are used; I have never ever found a single Japanese online store that accepts Ä/Ö in input (not even Amazon.co.jp). Add a name and address and you're suddenly give a "Please enter only English characters" error.
You'd think the Japanese of all people would know how annoying it is when systems accept a very limited amount of character sets.
It's kinda strange since Amazon.co.jp is really easy and handy for international shopping since you can even use it in English. They even precollect the Finnish VAT and DHL ships stuff from there super fast (package leaves Japan on a Friday, is at my door on Monday).
Failing to provide support or consideration to out-groups is a quite common feature of Japanese society. Difficulties arising from those situations tends to be blamed on those on the receiving end of them for being “other.”
Like many other commenters, I am also not sure why you’d want “katsu” transliterated to “cutlet”, but the author didn’t choose to have a tool to do this out of ignorance of the Japanese language.
Regarding the name, it's true that the word "cutlet" isn't typically an example of where you'd want to use the foreign spelling functionality. However it had several other nice points:
1. It follows the tradition of naming Japanese language tools after food the author likes, like MeCab, Sudachi, my own fugashi, etc.
2. It has a concrete image that's easy to convey, which is useful mnemonically
3. It's a very clear demonstration of the foreign spelling feature (even if not a useful one)
4. The word "katsu" is a good example for showing the difference between Hepburn and other romanization systems (katsu/katu)
"The cutlet was introduced to Japan during the Meiji period, in a Western cuisine restaurant in the fashionable Ginza district of Tokyo. The Japanese pronunciation of cutlet is katsuretsu.
In Japanese cuisine, katsuretsu or shorter katsu is actually the name for a Japanese version of the Wiener schnitzel, a breaded cutlet. Dishes with katsu include tonkatsu and katsudon."
If Cutlet had been around, this would have been exactly what I needed.
It might be easier for foreigner to understand the word, but at the same time no japanese will ever understand if you talk about cutlet curry instead of katsukaree. And that’s something i’d guess this would be used in; reading something out loud you don’t know how to read.
"Cutlet" strikes me as an edge case, considering that "katsu" has more or less become a word in its own right. So, perhaps an unfortunate choice of name for this project...
Translating debadora to "device driver" is the only way to make it understandable to someone who doesn't speak Japanese.
Then again, translating いとこ to "cousin" is also the only way to make that understandable to someone who doesn't speak Japanese (and does speak English), but the goal is to transliterate it as "itoko", not to translate it as "cousin". Like it or not, the Japanese word is "debadora".
The example above of “sale” is a good one. In a Japanese shop for Japanese and by Japanese, you will see “sale” but you will never see se-ru.
> No, the Japanese word is デバドラ.
False; these are one and the same claim, not two conflicting claims.
> Once you translate or transliterate it, you’ve diverged from the original
Again, this is wrong. If your "transliteration" has diverged from the original, it's not a transliteration. A transliteration is a reflection of the original, just using a different orthographic system.
The key factor you seem to be missing is that most of these words are themselves very lossy transliterations of original English words. Some become so ingrained in Japanese that they’re considered Japanese words by most people, and would make sense to leave in their lossy, katakana form even if romanized. As many people have mentioned here, “katsu” is ironically one of those words.
However, words like katsu are the exception and most are still considered English (or other language) words with the katakana being a best-effort representation in Japanese. The links shared by Zarel[0] and numpad0[1] demonstrate this perfectly.
The “original” word being expressed is “home” and “ホーム” is just the best you can do in an official Japanese script. “Ho-mu” isn’t useful for anyone, foreign or Japanese. It’s meaningless to non-Japanese speakers who aren’t familiar with the lossy transliteration to Japanese they need mentally reverse it. Japanese speakers could work it out quickly enough, but it would take more time and confusion since it’s neither representation they’re familiar with everyday, the original word “home” and the expected Japanese representation “ホーム.”
[0] https://magazine.jp.square-enix.com/biggangan/ [1] https://ja.wikipedia.org/wiki/タミル語
I don’t think it’s a settled question one way or the other, but is highly variable on context and history. For example, there have been several pushes, beginning with the Meiji to Revolution, to switch Japan to writing exclusively in romaji. If any had succeeded, we’d view this debate quite differently:
When an “alphabet” representation of a word is used in Japanese texts, the version of the word in either the original language or English is used.
e.g.
> タミル(Tamiḻ)という名称は、ドラミラ Dramiḻa(ドラヴィダ Dravida)の変化した形という説もある。[0]
https://magazine.jp.square-enix.com/biggangan/
For instance: notice that this says "big gangan" rather than "biggu gangan" (which is how you'd transliterate ビッグガンガン).
Also notice the other labels: "PRESENT", "twitter", "BIGGANGAN OKAWARI".
Perhaps it's hard to say much more without knowing who these transliterations are supposed to serve. If it's a general Japanese-speaking population with limited knowledge of English, I'm not sure whether they'd prefer romaji or the original spelling.
Ex., to how many people is it useful to convert アルバイト from "arubaito" to "Arbeit", especially when the Japanese word has a different connotation to the German (part-time work vs occupation)
Real world example[0]
I myself am not too sure what to think about カツ becoming Cutlet when カツ is a shortened version of カツレツ, which itself is the katakana version of Cutlet. Technically, カツ would be Cut, but then English doesn't have that notion of using Cut to mean Cutlet.
There are plenty real world examples for “Beef curry” and “ビーフカレー” used in pairs in Japanese, katsu is usually an extra so it’s rarer though.
They have four different systems. There is Kangi which is from Chinese. There are two different systems that are phonetic but based on syllables rather than single sounds. And finally they use the Roman script like we do. It must be really hard on the elementary school kids who are trying to learn this all.
I'm also pretty unconvinced of the article's stated purpose of "readable" URL text. I spot checked a few Japanese news sites (Yomiuri, Asahi, Mainichi), and none of them try to do this. The URLs just have random alphanumeric article IDs. I don't think romaji is valuable for most readers, who probably find it more effort to decipher it than to simply read the Japanese text in the title.
Here's an example:
https://magazine.jp.square-enix.com/biggangan/
(Notice: "big gangan" rather than "biggu gangan".)
While it's true that Japan tends to use numbers for dynamic content URLs, this is more about (and using CMSes that require them) than users actually preferring them.
Japanese URLs do frequently tend to be in English, though.
Your example actually requires human decision-making, since Big Gangan is a proper noun with a canonical English name. I expect most Japanese devs to be as familiar with romanization as English devs are with spelling and grammar, so a tool shouldn't be needed unless you're dealing with a large amount of text and can allow for errors.
If you're comfortable with Python you should be able to use fugashi to throw together a tool pretty quickly, I've though about doing it before but never got around to it.
What Python would look like in Japanese:
https://www.reddit.com/r/ProgrammingLanguages/comments/g9iu8...
For web dev, I think it's probably fine.