These are less JavaScript problems than utf-16 problems. The whole one character is not a code point problem. It's common to Java, .net, basically all of windows, and anything else that uses utf-16 strings. The solution is easy. If you need a one to one mapping of code points to characters convert to utf32 first. Utf8 has the same problems, the only difference is people know characters and code points don't match up. Whereas with utf16 there's a bunch of people who are either new or should never have been programmers to begin with that are clueless about it. Sadly this number is so large that just about any program that uses utf-16 strings is broken for inputs where code points != characters. This is partly the fault of the languages and libraries which give you functions like substring, reverse, etc on utf-16 strings, where they basically have no consistent meaning. It should have been a storage format not a manipulation format.