How to chop off bytes of an UTF-8 string to fit into a small slot and look nice
domm.plix.at
domm.plix.at
* https://hoytech.github.io/truncate-presentation/ / https://metacpan.org/pod/Unicode::Truncate
* https://unicode.org/reports/tr29/#Grapheme_Cluster_Boundarie...
Truncating at codepoint boundaries at least avoids generating invalid (non-UTF-8) strings, but can still result in confusing or incorrect displays for human readers, so for best results the truncation algorithm should take extended grapheme clusters into account, which are probably the closest thing that Unicode has to what most people think of as "characters".
EDIT: oh lord Hacker News strips emoji. You get the idea even though HN ruined the illustrations. Not my fault HN is broken.
Just because the worst case can't improve doesn't mean that making the average case better is worthless.
(I don't think the talk where the stuff about Arabic was mentioned was Plain Text by Dylan Beattie, but I haven't re-watched it to confirm. So maybe it is. Can't remember the name of any other talk about the subject right now.)
IIRC the lower levels of Windows will happily work with filenames that are not valid Unicode strings, for example if you use the kernel API rather than Win32.
But what about Win32? If you create a file before normalization and then open it using the normalized form, will it open the same file or return file not found?
What about other systems? For example AWS' S3 allows UTF-8 keys, with no mention of normalization[1].
On the phone so can't try myself right now.
Anyway for general text I agree, but for identifiers, filenames and such I prefer to treat them as opaquely as possible.
[1]: https://docs.aws.amazon.com/AmazonS3/latest/userguide/object...
Just made two files named ä.txt which happily sat next to each other, one being NFC and other NFD.
So yeah, don't mess with the normalization of filenames.
I don't use Windows, so I can't check. Linux literally allows any arbitrary byte except for 0x00 and 0x2F ('/' in ASCII/UTF-8). It's a problem for programming languages that want to only use valid Unicode strings, like Python. Rust has a separate type "OsString" to handle that, with either lossy conversion to "String" or a conversion method that can fail. Python uses the custom use Unicode range to represent invalid byte sequences in filenames. It's all a mess. JavaScript doesn't give a damn about the validity of their UTF-16 strings.
(Note that Rust's OsString is different from it's CString type. Well, I guess under Unix they're the same, but under Windows OsString is UTF-16 (or "WTF-16", because it isn't actually valid UTF-16 in all cases).)
I tried using U+13161 EGYPTIAN HIEROGLYPH G029[1], which resulted in a string of length 2 as expected.
Using both chars (code units) and just the first char (code unit) worked equally fine. In Windows Explorer the first one shows the stork as expected, while the second shows that "invalid character" rectangle.
So yeah, treating filenames as nearly-opaque byte sequences is probably the best approach.
Äh, nu går vi och gör något annat.
Win32/NTFS treats them as two different filenames, as it doesn't normalize them before storing/comparing.
The Linux kernel doesn't validate filenames in any way, so a filename in Linux can contain any byte except 0x2F ('/', which is interpreted as directory separator) and 0x00 (which signals the end of the byte string).
ETA: of course some file systems have other limitations, for example '\' is not valid in FAT32.
import regex as re
def grapheme_clusters(text):
# \X is the regex pattern that matches a grapheme cluster
pattern = re.compile(r'\X')
return [
match.group(0)
for match in
pattern.finditer(text)
]
https://chatgpt.com/share/481c9c94-0431-4fcb-82aa-a44a4f3c21...This article is very specific to Perl, and the way it does so is also subject to question - it does not look efficient.
You will be better off by reading excellent wikipedia page on UTF-8: https://en.wikipedia.org/wiki/UTF-8
Now, extended grapheme cluster enumeration is much more complex than finding the next non-continuation byte (or counting such), but to perform those correctly you would ultimately end up reading the official spec at unicode.org and perusing reference implementations like ICU (which is painful to read) or from standard library/popular packages for Rust/Java/C#/Swift (the decent ones I'm aware of, do not look at C++).
-"Don't do this. Instead use a completely different programming language."
https://github.com/toml-lang/toml/issues/994
https://github.com/toml-lang/toml/issues/966
Great to see new faces in the community, sad to see the sheer insanity of MARC21 still causing chaos. MARCXML is gonna make it obsolete Any Day Now!
BIBFRAME is gonna make it obsolete Any Day Now!
("any day now" in this sector means that librarians have been talking about it for two decades and in about two decades something might actually happen)
The reason libraries don't do cataloguing any longer anywhere near as much hasn't got much to do with MARC 21 being hard.
Having catalogers or metadata librarians write directly in MODS XML by hand never made sense (although some folks tried this), but as far as something usable to ship around I'd rather get MODS than MARC or dublin core. I really don't want to have to query a triple store to aggregate records.
Catalogers ideally would have tools that make it easy for them follow RDA/AACR2 descriptive practices without having to think about the details of MARC or MODS or linked data.
I've been out of the business for a couple of years, so I have not been following Library of Congress' BIBFRAME transition.
(*Rendered version on Wikipedia: https://commons.wikimedia.org/wiki/File:Lateef_unicode_U%2BF... )
> The real problem is that USMARC uses an int with 4 digits to store the size of a field, followed by 5 digits for the offset.
A colleague told me they used to exploit this "feature" to leave hidden messages in MARC records between fields.
Hidden is a very relative term there.