Finding CSV files that start with a BOM using ripgrep
til.simonwillison.net
til.simonwillison.net
That's not correct because the `m` flag gets enabled by the multiline option.
$ printf 'a\nbaz\nabc\n' | rg -U '^b'
baz
Need to use `\A` to match start of file or disable `m` flag using `(?-m)`, but seems like there's some sort of bug though (will file an issue soon): $ printf 'a\nbaz\nabc\n' | rg -U '\Ab'
baz
$ printf 'a1\nbaz\nabc\n' | rg -U '\Ab'
baz
$ printf 'a12\nbaz\nabc\n' | rg -U '\Ab'
$The bug is fixed on master. Thanks for calling this to my attention! https://github.com/BurntSushi/ripgrep/issues/1878
1. These seem like they're effectively integration tests. They check the entire ripgrep command line app works as intended. Is this because the bug was not where it looks like it is, in the regex crate, but elsewhere? If not, it seems like they'd be better as unit tests closer to where the bugs they're likely to detect would lie?
2. While repeating yourself exactly once isn't automatically a bad sign, it smells suspicious. It seems like there would be a lot of tests that ought to behave exactly the same with or without --mmap and so maybe that's a pattern worth extracting.
> While repeating yourself exactly once isn't automatically a bad sign, it smells suspicious. It seems like there would be a lot of tests that ought to behave exactly the same with or without --mmap and so maybe that's a pattern worth extracting.
Quite possibly. Not a bad idea. I already do that in the tests for PCRE2. Most tests are run with both the default regex engine and with PCRE2. (There are some tests that have intended behavioral differences with mmap enabled, but those can be handled on a case-by-case basis.)
Is this documented? RG(1) only says
-m, --max-count <NUM>
Limit the number of matching lines per files searched to NUM.
Is this a different option or is there an implied <NUM> that prevents 'rg -U ^' from searching from the beginning of the file?ripgrep also has a -U/--multiline flag, but it's orthogonal to the regex mode called multiline. It's an unfortunate naming clash, but they are names in otherwise distinct namespaces.
ripgrep always enables the regex flag 'm', regardless of whether -U/--multiline is enabled or not.
I initially thought it was searching for "Bill of Materials" for electronics projects or similar.
https://en.wikipedia.org/wiki/Byte_order_mark#UTF-8
https://www.w3.org/International/questions/qa-utf8-bom.en.ht...
> In UTF-8, the BOM corresponds to the byte sequence <EF BB BF>. Although there are never any questions of byte order with UTF-8 text, this sequence can serve as signature for UTF-8 encoded text where the character set is unmarked. As with a BOM in UTF-16, this sequence of bytes will be extremely rare at the beginning of text files in other character encodings. For example, in systems that employ Microsoft Windows ANSI Code Page1252, <EF16 BB16 BF16> corresponds to the sequence <i diaeresis, guillemet, inverted question mark> “ï » ¿”.
In practice, the UTF-8 BOM pops up. I usually see it on Windows.
[1] - http://www.unicode.org/versions/Unicode13.0.0/ch23.pdf#G1963...
That was a fun, and totally unstressful way to begin my time managing a racing league's race events.
Well require is a bit excessive, but it certainly allows and recommends one.
Utf8 does not need one because the code units are bytes, so bytes order is not a concern.
Exchanging utf32 is pretty rare though, and as long as you don’t move anything between machines bytes order is not an issue.
printf '\xEF\xBB\xBF' >bom.dat
find . -name '*.csv' \
-exec sh -c 'head --bytes 3 {} | cmp --quiet - bom.dat' \; \
-print
The -exec option for find can be used as a filter (though -exec disables the default action, -print, so it must be reenabled after).Could be made into a oneliner by replacing the 'bom.dat' argument to cmp with '<(printf ...)'.
-n, --bytes=LIMIT compare at most LIMIT bytes
so head is not really necessary: find . -name '*.csv' -type f -exec cmp -sn 3 {} bom.dat \; -print
Using -exec as a filter is a nice feature more people should use. That -type was put there just to avoid directories.Besides, at the beginning people were really against variable size encodings. UTF-8 won despite the Unicode consortium and all the committees effort, not because of it.
UTF-16 is not a fixed-size encoding thanks to surrogate pairs. UCS-2 is a fixed-size encoding but can’t represent code points outside the BMP (such as emoji) which makes it unsuitable for many applications.
Besides, most of the time individual code points aren’t what you care about anyway, so the cost of a variable-sized encoding like UTF-8 is only a small part of the overall work you need to support international text.
* all content will increase significantly in size using utf32 (utf16 is also variable-size, and markups being extremely common and usually ascii, while utf8 is not a guaranteed winner against 16 it often is)
* unicode itself is variable-size encoding due to combining code points, so a fixed-size encoding really doesn’t net you anything
That sound more like an anti-feature resulting in unstable programs that almost work.
Either way, a complete rewrite of the text handling functionality should give you flawless functionality. At this point in time, all the important ecosystems that use UTF-8 are almost there.
This is very different from the other encodings where a complete rewrite of the text handling functionality is needed just not to fail every time. That made all the important ecosystems that used other encodings to get almost there much sooner, but there was an important period when everything was broken, and the improvements are much slower nowadays, because when you need to fix every aspect of something, iterations take much more labor.
It's both. It will generally ignore non-ascii data, but that is very commonly something you don't care about, in which case it's a net advantage over plain not working at all.
0: Similarly, U+1F1 "DZ" is two characters, but one Unicode code point, which is much, much worse as it means you can no longer treat encoded strings as concatenations of encoded characters. UTF-8-as-such doesn't have this problem - any 'string' of code points can only be encoded as the concatenation of the encodings of its elements - but UTF-8 in practice does inherit the character-level version of this problem from Unicode.
I base this number off the "Stream-Safe Text Format" which suggests that while it's preferred that you accept infinitely-long characters, a cap of 31 code points is more or less acceptable.
However, as you say, by 1996 people were already using the older UCS-2 standard.
And sure they updated it in 2003 but "don't use invalid codepoints" is not a really notable update.
Unicode / ISO 10646 is specifically defined to only have code points from 0 to 0x10FFFF. As a result UTF-8 that would decode outside that range is just invalid, no different from if it was 0xFF bytes or something.
It also doesn't make sense to write UTF-8 that decodes as U+D800 through U+DFFF since although these code points exist, the standard specifically reserves them to make UTF-16 work, and you're not using UTF-16.
You can't tell me what to do, dad. I'll encode 64 bits and you can't stop me! Bwahahahaa!
$ perl -MEncode=encode_utf8 -e'print encode_utf8 "\x{7fff_ffff_ffff_ffff}"' | hex
0000 ff 80 87 bf bf bf bf bf bf bf bf bf bf ÿ␀␇¿¿¿¿¿¿¿¿¿¿Furthermore, even if you assume a implied zero bit at position -1, that would only be FF BF BF BF BF BF BF BF, with value U+3FF'FFFF'FFFF.
Also 7FFF'FFFF'FFFF'FFFF is only 63 bits - fer chrissakes son, learn to count.
That's needlessly pedantic. If you use an old version of the spec those bytes are valid.
And "have the capability" seems to me to be talking about what the underlying method is able to do, not the full set of "must not" rules.
Code points might be combined to form graphemes and grapheme clusters. Some of the latest emojis are extended grapheme clusters, for e.g. handling the combinatorics of mixed families. This is a higher level composition than UTF-x, it's logically a separate layer.
IMO talking about characters in the context of Unicode is often unhelpful because it's vague.
In programming languages or APIs where precision matters, your goal should be to avoid this notion of characters as much as practical. In a high level language with types, just do not offer a built-in "char" data type. Sub-strings are all anybody in a high level language actually needs to get their job done, "A" is a perfectly good sub-string of "CAT" there's no need to pretend you can slice strings up into "characters" like 'A' that have any distinct properties worth inventing a whole datatype.
If you're writing device drivers, once again, what do you care about "characters"? You want a byte data type, most likely, some address types, that sort of thing, but who wants a "character" ? However, somewhere down in the guts a low-level language will need to think about Unicode encoding, and so eventually they do need a datatype for that when a 32-bit integer doesn't really cut it. I think Rust's "char" is a little bit too prominent for example, it needn't be more in your face than say, std::num::NonZeroUsize. Most people won't need it, most of the time and that's as it should be.
31.