What Every Software Developer Must Know About Unicode and Character Sets (2003)
joelonsoftware.com
joelonsoftware.com
What libraries do C and C++ programmers use to hold unicode strings and convert between encodings these days?
http://en.wikipedia.org/wiki/C%2B%2B0x#New_string_literals
There's also a link to a proposed boost solution here:
http://stackoverflow.com/questions/511280/is-there-stl-and-u...
Doesn't quite sound like the standard way that you're looking for, but moving closer.
Edit:
"ICU today is de-facto standard Unicode/localization library" from a mailing discussion of the boost solution. And http://art-blog.no-ip.info/cppcms/blog/post/43 has an interesting comparison of a few libraries, but not too comprehensive.
http://diveintopython3.org/strings.html
"Some Boring Stuff You Need To Understand Before You Can Dive In."
Do I copy the code points, or the encoded characters--the bytes--along with what encoding is used? Similarly, when I paste, is it the code points I paste which are instantaneously encoded using the application's encoding scheme?
In PHP curl_exec returns data in the raw encoding of the source. Fine, some people will want that. But I want to do things with the data, so I want it in UTF8.
So, I ended up writing my own curl_exec_utf8 function which I'm sure is wrong for many edge cases, but it is 2010! Why is there no decent ways to deal with charsets?
Here is the function, in case any of you need it, or want to point out how it is hopelessly broken : http://stackoverflow.com/questions/2510868/php-convert-web-p...
What he says is still very much valid. It's a nice intro into Unicode. Quite refreshing to read.
If it would overflow, you just set the long to the max size, byte the bullet and read the entire string. If you are using a dynamic language just throw whatever InfinitelyLargeNumber class it has in the size column and you're good.
If you're worried about ram that much you can just use plain strings when you need to.
Why not link to that in the first place?
You have to do that with all Unicode encodings, since a semantic symbol can be composed of multiple combining characters.