Strings, bytes, runes and characters in Go
blog.golang.org
blog.golang.org
Other languages like C# seem to be different on the surface, but in fact they index and measure by code units as well (2 byte UTF-16 code units), not by code points.
You can't usefully index a unicode stream in constant time and do correct and useful textual stuff anyway due to combining codepoints which may not have precombined forms (if only because there is no defined limit to the number of combining codepoints tacked onto the base) (so normalization will not save you) or codepoints which are not visible to the user and which you may or may not want to see depending on the work you're doing.
People really need to come to terms that a unicode stream is exactly that, a stream.
To find an index of a substring you need to scan the string, right. But once you have the byte index you can quickly jump to its position in the string, e.g. when you do a slice operation based on that index: s[i:]. If strings.Index() returned a code point index and not a byte index you would have to scan the string again.
Stop doing that and just get the bit of string you want in the first place?
How about indexing, measuring and slicing operations based on user-perceived characters, then?
There is a difference between a string and a byte array. A string is a string. A byte array is a []byte (byte slice). You have to explicitly cast from one to another. Neither are inherently utf8. A string is represented by a byte array under-the-hood, and string literals in your code are read as utf8 encoded. Strings themselves are not necessarily utf8 encoded, and if you need to use a different encoding there's libraries for that (unless you're using something really esoteric).
The distinction between a string with one encoding and a string with another is subtle but vitally important - exactly the sort of thing a type system should take care of.
If you're expecting to get data in other other encodings you could put together some detection and transformation at the point of ingress and convert to UTF-8 encoded text for the rest of your application.
The `next(string, index)` function used for the iteration protocol works like the `utf8.DecodeRuneInString()` shown in the example, but it returns the next valid index rather than the character width.
(check the docs, there's also RuneStart, which returns true/false if you index mid-rune)
In C# you can encode/decode strings to byte arrays based on your desired encoding, but a string is composed of characters, it's in memory representation is abstracted.
Is this a performance or zero copy thing? Not having to encode/decode to get to the bytes?
Microsoft had reasons to pick UTF-16 for C# and the CLR. The Windows API speaks UTF-16, and when C# was first announced way back in 2000, UTF-8 was not yet a widely used encoding on the Web; Unicode was sadly not that widely used on the Web, period. The decision to use 16-bit Unicode in the Win32 API went even further back, to the development leading up to NT 3.1's release in July 1993. At that point, UTF-8 was a relative baby; it was presented at USENIX in January 1993. Also, back then, code points basically were 16 bits because surrogate pairs were but a twinkle in the Unicode Consortium's eye. The Unicode 2.0 standard added surrogate pairs to help them expand their CJK selection and generally let them add more chars of all sorts; it wasn't released until 1996. Now here we are and we have 😃, U+1F603.
Go came along in 2009; by then, many Web sites were being served in UTF-8, and Go'd be used in significant part in Web operations, and UTF-8 was also the default encoding in many Unix environments. Go initially didn't run on Windows at all and was ported by the community, so fitting in with Win32 wasn't an issue. Code points that wouldn't fit in 16 bits were a fact of life by '09, too. Arguments about inherent merits of encodings aside, it probably would have seemed to lots of folks that UTF-8 was a natural choice for that task at that time.
Also, Pike and Thompson are two of the three co-designers of Go and co-designed UTF-8, and UTF-8 was the encoding used by the Plan 9 OS/environment they built at Bell Labs, so, again, technical details aside, it was kind of a foregone conclusion which encoding they'd build the language around. :)
The one thing I do not want to do here is get in an argument about the inherent merits of character encodings, so I'm just not gonna do that. :)
On UTF-16 and Windows NT: http://support.microsoft.com/kb/99884, http://en.wikipedia.org/wiki/Windows_NT_3.1, and http://en.wikipedia.org/wiki/Unicode
On Unicode adoption on the Web, UTF-8, and Plan 9: http://googleblog.blogspot.com/2012/02/unicode-over-60-perce..., http://en.wikipedia.org/wiki/UTF-8, http://en.wikipedia.org/wiki/Plan_9_from_Bell_Labs
Perhaps I should have left out the word "internal" since it's exposed.
That's a hard problem, and avoiding it in every situation would require scanning the strings for surrogates beforehand, when you might never need to know that information. Go makes it explicit that knowing the exact character position and string length in characters comes at a cost.
There's a good discussion of this on Tim Bray's blog: http://www.tbray.org/ongoing/When/200x/2003/04/26/UTF
You can even export them without Capital letters ;-)
// Check that de ≡ 1 mod p-1, for each prime.
// This implies that e is coprime to each p-1 as e has a multiplicative
// inverse. Therefore e is coprime to lcm(p-1,q-1,r-1,...) =
// exponent(ℤ/nℤ). It also implies that a^de ≡ a mod p as a^(p-1) ≡ 1
// mod p. Thus a^de ≡ a mod n for all a coprime to n, as required.
Sadly, the spec requires identifiers to be just Unicode letters and digits, so we will never experience the power and glory of emoji function names in Go.Note that many languages that appear to offer indexing by rune (e.g., Java) do not in fact do so, since their 16-bit "character" type is incapable of representing all runes. The fact that this is only rarely an issue points at the fundamental rarity with with code needs to deal with runes-qua-runes.