Java loaded full unicode code point semantics into its standard `java.lang.String` class. These _are not guaranteed_ to have `O(1)` performance characteristics, because the underlying storage format is dynamically either a UTF-16-esque variant (with surrogate pairs for characters that don't fit in 16 bit), or a single-byte-per-char format if the string does not contain any non-ASCII. This has the advantage of being very very slightly more obvious, given that both methods exist and are documented:
void main() {
String x = "(that emoji here)";
System.out.println("Chars: " + x.length());
System.out.println("Codepoints: " + x.codePointCount(0, x.length()));
System.out.println("As stream of chars (= UTF16-esque with surrogate pairs):");
x.chars().forEach(System.out::println);
System.out.println("As a stream of codepoints:");
x.codePoints().forEach(System.out::println);
}
This ends up printing: Chars: 7
Codepoints: 5
As stream of chars (= UTF16-esque with surrogate pairs):
55358
56614
55356
57340
8205
9794
65039
As a stream of codepoints:
129318
127996
8205
9794
65039
NB: Apparently many hackernews readers know java but don't use it all that often day-to-day. The provided java snippet is vanilla valid and can be executed with `java ThatFile.java` (no need to compile it first), though it does use preview features.The fact that the codepoint counter is a very awkward `codePointCount` call has the dubious benefit of highlighting this method loops through and therefore would be quite slow on very large strings.