> However, this solution does not work very well for the Japanese community. For a variety of complicated reasons, Japanese encoding, such as SHIFT-JIS, are not considered to losslessly encode into UTF-8. As a result, Ruby has a policy of not attempting to simply encode any inbound String into UTF-8.
> This decision is debatable, but the fact is that if Ruby transparently transcoded all content into UTF-8, a large portion of the Ruby community would see invisible lossy changes to their content. That part of the community is willing to put up with incompatible encoding exceptions because properly handling the encodings they regularly deal with is a somewhat manual process.
Does someone has any sources about this?
Esp, why is SHIFT-JIS important? Is Unicode not capable of encoding the full Japanese language? What information would you loose by encoding with Unicode?
And why can't Unicode be extended to encode this missing information?
This seems to be a much better solution than this encoding nightmare. Esp in a language like Ruby, I want to have it simple and straightforward.