Regex character "$" doesn't mean "end-of-string"
sethmlarson.dev
sethmlarson.dev
Huh. I always think of them as "start-of-line" and "end-of-line". I mean, a lot of the time when I'm working with regexes, I'm working with text a line at a time so the effect is the same, but that doesn't change how I think of those operators.
Maybe because a fair amount of the work I do with regexes (and, probably, how I was introduced to them) is via `grep`, so I'm often thinking of the inputs as "lines" rather than "strings"?
$ printf 'Line with EOL\nLine without EOL' | grep 'EOL$'
Line with EOL
Line without EOL
$ grep --version | head -n1
grep (GNU grep) 3.8It's not matching the newline character after all.
man 7 regex: '$' (matching the null string at the end of a line)
pcre2pattern: The circumflex and dollar metacharacters are zero-width assertions. That is, they test for a particular condition being true without consuming any characters from the subject string. These two metacharacters are concerned with matching the starts and ends of lines. ... The dollar character is an assertion that is true only if the current matching point is at the end of the subject string, or immediately before a newline at the end of the string (by default), unless PCRE2_NOTEOL is set. Note, however, that it does not actually match the newline. Dollar need not be the last character of the pattern if a number of alternatives are involved, but it should be the last item in any branch in which it appears. Dollar has no special meaning in a character class.
Vim is what did that for me.
> Matches the start of the string, and in MULTILINE mode also matches immediately after each newline.
Or JavaScript:
> An input boundary is the start or end of the string; or, if the m flag is set, the start or end of a line.
\A and \Z are start/end of input regardless of mode… when they’re available, that’s not the case of all engines.
No? It’s a semantics decision.
Usually ^ matches only at the beginning of the string, and $ matches only at the end of the string and immediately before the newline (if any) at the end of the string. When this flag is specified, ^ matches at the beginning of the string and at the beginning of each line within the string, immediately following each newline. Similarly, the $ metacharacter matches either at the end of the string and at the end of each line (immediately preceding each newline).
In single-line [2] mode, the line starts at the start of the string and ends at the end of the line where the end of the line is either the end of the string if there is no terminating newline or just before the final newline if there is a terminating newline.
In multi-line mode a new line starts at the start of the string and after each newline and ends before each newline or at the end of the string if the last line has no terminating newline.
The confusion is that people think that they are in string-mode if they are not in multi-line mode but they are not, they are in single-line mode, ^ and $ still use the semantics of lines and a terminating newline, if present, is still not part of the content of the line.
With \n\n\n in single-line mode the non-greedy ^(\n+?)$ will capture only two of the newlines, the third one will be eaten by the $. If you make it greedy ^(\n+)$ will capture all three newlines. So arguably the implementations that do not match cat\n with cat$ are the broken ones.
[1] https://docs.python.org/3/howto/regex.html#more-metacharacte...
[2] I am using single-line to mean not multi-line for convenience even though single-line already has a different meaning.
You seem to have redefined “line” as “not a line”.
> The confusion
I’m sure redefining “line” as “nothing like what anyone reasonable would interpret as a line” will help a lot and right clear up the confusion.
Works for me.
How do you square that with your assertion that in your invention of "single-line mode" you implicitly define "line" as matching \n\n?
I suspect the real reason for Python's behavior starts with the early decision to include the terminating newline in the string returned by IOBase.readline().
Python's peculiar choice has some minor advantages: you can distinguish between files that do and don't end with a terminating newline (the latter are invalid according to POSIX, but common in practice, especially on Windows), and you can reconstruct the original file by simply concatenating the line strings, which is occasionally useful.
The downside of this choice is that as a caller you have to deal with strings that may-or-may-not contain a terminating newline character, which is annoying (I often end up calling rstrip() or strip() on every line returned by readline(), just to get rid of the newlines; read().splitlines() is an option too if you don't mind reading the entire file into memory upfront).
My guess is that Python's behavior is just a hack to make re.match() easier to use with readline(), rather than based on any principled belief about what lines are.
The very post we're commenting on shows that that's not true: PHP, Python, Java and .NET (C#) share one behavior (accept "\n" as "$"), and ECMAScript (Javascript), Golang, and Rust share another behavior (do not accept "\n" as $).
Let's not argue about which is “the most common”; all of these languages are sufficiently common to say that there is no single common behavior.
> $ matches at the end of the string or before the last character if that is a newline, which is logically the same as the end of a single line.
Yes, that is Python's behavior (and PHP's, Java's, etc.). You're just describing it; not motivating why it has to work that way or why it's more correct than the obvious alternative of only matching the end of the string.
Subjectively, I find it odd that /^cat$/ matches not just the obvious string "cat" but also the string "cat\n". And I think historically, it didn't. I tried several common tools that predate Python:
- awk 'BEGIN { print ("cat\n" ~ /^cat$/) }' prints 0
- in GNU ed, /^M/ does not match any lines
- in vim, /^M/ does not match any lines
- sed -n '/\n/p' does not print any lines
- grep -P '\n' does not match any lines
- (I wanted to try `grep -E` too but I don't know how to escape a newline)
- perl -e 'print ("cat\n" =~ /^cat$/)' prints 1
So the consensus seems to be that the classic UNIX line-based tools match the regex against the line excluding the newline terminator (which makes sense since it isn't part of the content of that line) and therefore $ only needs to match the end of the string.The odd one out is Perl: it seems to have introduced the idea that $ can match a newline at the end of the string, probably for similar reasons as Python. All of this suggests to me that allowing $ to match both "\n" and "" at the end of the string was a hack designed to make it easier to deal with strings without control characters and string that end with a single newline.
But /\n/ does
If you read a line, you usually remove the newline at the end but you could also keep it as Python does. If you remove the newline, then a line can never contain a newline, the case cat\n can never occur. If you keep the newline, there will be exactly one newline as the last character and you arguably want cat$ to match cat\n because that newline is the end of the line but not part of the content. It makes perfect sense that $ matches at the end of the string or before a newline as the last character as it will do the right thing whether or not you strip the newline.
If you want cat$ to not match cat\n, then you are obviously not dealing with lines, you have a string with a newline at the end but you consider this newline part of the content instead of terminating the line. But ^ and $ are made for lines, so they do not work as expected. I also get what people are complaining about, if you are not in multi-line and have a proper line with at most one newline at the end, then it will behave exactly as if you are in multi-line which raises the question why you would have those two modes to begin with. Not multi-line only behaves differently if you have additional newlines or one newline not at the end, that is if you do not have a proper line, so why should $ still behave as if you were dealing with a line?
If you have a file containing `A\nB\nC` in a file, the file is three lines long.
I guess it could be argued that a file containing `A\nB\nC\n` has four lines, with the fourth having zero length.
That a regex is applying to an in memory string vs a file doesn't feel to me like it should have different semantics.
Digging into the history a little, it looks like regexes were popularized in text editors and other file oriented tooling. In those contexts I imagine it would be far more common to want to discard or ignore the trailing zero length line than to process it like every other line in a file.
https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1... ("text file")
It's a file with zero complete lines. But it has 1 line, that's incomplete, right?
The file starts empty. Anything in it starts "a line". So it's 1 incomplete line.
I hate weird states.
Semantically many libraries treat that as a line because while \n<EOF> means "the end of the last line" having just <EOF> adds additional complexity the user has to handle to read the remaining input. But by the book it's not "a line".
If I said "ten buckets of water" does that mean ten full buckets? Or does a bucket with a drop in it count as "a bucket of water?" If I asked for ten buckets of water and you brought me nine and one half-full, is that acceptable? What about ten half-full buckets?
A line ends in a newline. A file with no newlines in it has no lines.
Assuming that EOF is identical to \\nEOF will end up causing trouble for you one day, because it's not actually identical.
Technically as per posix a file as you describe is actually a binary file without any lines. Basically just random binary data that happens to kind of look like a line.
In practice, most utilities expecting text files will still operate on it.
"Lines" are a convention established by (or not) software reading a data stream.
Also, why couldn't you have a text file without any lines?
It's a file with zero complete lines. But it has 1 line, that's incomplete, right?
Because the Unix definition of text file requires the file to end with a newline. "Lines" only exist in the context of text files. If there's no terminating newline, it's (pedantically) not a text file and so has no lines. Now, in practice, if you open() that file in text mode, it doesn't TMK return an error if the terminating newline isn't present, but it's undefined behaviour.And if you do have a terminating newline, then you have at least one line :).
This isn't a weird state. It's a language problem. An 'incomplete line' isn't a type of line, it's an unfortunate name for a thing that is not a line. Just like how the 'wor' is an incomplete word (the word 'word'), but 'wor' is, of course, not a word.
Same thing for formalisms like equations in algebra or formulas in propositional logic— we have the phrase 'well-formed formula', and we might describe some sequences of terms as 'incomplete formulas' or perhaps 'ill-formed formulas', but those phrases don't describe anything that meets the formal system's definition of 'formula' at all— they are not formulas. 'Ill-formed formula' is not a compositional phrase where 'ill-formed' describes a feature of a 'formula'. It's a bit of convenient language for what we can intuitively or metaphorically recognize as a formula-ish thing.
A sequence of zero or more non- <newline> characters plus a terminating <newline> character.
See also ‘3.403 Text File’ for the definition of a text file. No new line characters, no lines. No lines, not a text file.
That seems like a broken (maybe just bad?) definition/specification to me. A blob of JSON in a file isn't "text" if there's no newline character trailing it?
As a spec it’s fine. It defines a text file in such a way that you can easily write code to process such a file deterministicaly.
$ echo -n "A" | wc --lines
0https://stackoverflow.com/questions/729692/why-should-text-f...
$ echo -n foo | wc -l
0 foo\nbar\n
Would be 2 lines in *nix and 3 lines in windows.This is definitely not one of them.
You don't have to love a company to acknowledge they did something right.
To indicate that a serially received line is complete, the interpretation as a terminator makes perfect sense - abcd\n is a complete line, abc is a still incomplete line. In a text file the interpretation as a separator might be preferable because that gets rid of the issue of the last line not having a newline - a\nb\nc are three lines separated by two newlines, a\nb\nc\n are four lines separated by three newlines and the last line is empty.
But then it might also be useful to have a terminator in a text file to be able to detect an incompletely written line. So using two characters, one for each purpose, could solve the problem. \r means the line is complete, \n means it follows a next line. abc is an incomplete line, abcd\r is a complete line and no line follows, abcd\r\n is a complete line and a second incomplete line follows which is currently empty. abcd\r\n\r are two complete lines, the second one empty. abcd\r\nefg is a complete line followed by an incomplete line. abcd\r\nefg\r are two complete lines. You could even have two incomplete lines abc\nefg.
But I think Windows always uses \r\n because this is how you get to a newline on a typewriter or really old printer, you return the carriage and feed the paper one line. I do not think that they had the idea of differentiating between terminator and separator, otherwise you could have only \r and maybe even only \n sometimes. But in principle this could work quite nicely, I guess. You could start a line with \n and end it with \r, this would give you \r\n between lines and \r after the final line. Or nothing if the final line is incomplete or \r\n if the final line is incomplete and currently empty. The odd thing would be a newline as the very first character, maybe one could suppress that. This would also be compatible with Windows and nix, it would just consider all nix lines incomplete. Only abc\rdef\r would not really make sense, two complete lines but the second one is not a new line.
If I ever get to write a new operating system, I will inflict this on humanity.
Very very technically a "newline" character indicates the start of a new line, which is why it is not called the "end-of-line" character.
3.206 Line A sequence of zero or more non- <newline> characters plus a terminating <newline> character.
https://pubs.opengroup.org/onlinepubs/9699919799.2018edition...
getline() reads an entire line from stream, storing the address
of the buffer containing the text into *lineptr. The buffer is
null-terminated and includes the newline character, if one was
found.
...
... a delimiter character is not added if one was
not present in the input before end of file was reached.
EOF seems same as end-of-string.If a line is missing a newline then we just disregard it?!
A way better way to deal with newline is it's a separator like comma. And like in modern languages we allow a final separator, but ignore it so that is easier for tools to generate files.
Now all combinations of characters, including newline characters, has an interpretation without dropping anything.
But if you look beyond files, the interpretation as a terminator also makes perfect sense, when you receive text over a serial connection it signals that the line is complete which does not necessarily imply that another line will follow. The same in a file, if the terminating newline is missing, you can deduce that an incomplete write occurred and some data might be missing. If you decide to have a newline as a separator after the last line but to ignore it, then you can not represent an empty last line.
I guess you would need two different characters, one terminator and one separator. You could start a line with \n and end it with \r. The \n separates the line from the one before, then \r terminates the line and marks it as complete. You would get \r\n between lines as on Windows and the last line would only have \r if complete or would otherwise count as incomplete. Then again you could almost get the same thing with \n only, you would just have to change the interpretation, instead of \n giving you a line and no \n giving you not a line, you would have to say that \n gives you a complete line and no \n gives you an incomplete line. With that you could however not have an incomplete empty line.
A file that contains characters organized into zero or more lines [so characters with no newlines are OK]
No NUL, and lines (delimited by and including newline) not exceeding LINE_MAX bytes.
(“But backtick is annoying to type” said the Europeans.)
I would also imagine that, if this became the norm, IDEs would quickly standardize around common notation - probably actually based on existing regex symbols and escapes - to quickly input that, similar to TeX-like notation for inputting math. So if you're inside a regex literal, you'd type, say, \A, and the editor itself would automatically replace it with the Unicode sigil for beginning-of-string.
Certainly making the perlre library available separate to perl encouraged its widespread use, and lots of others copied or were inspired by it.
"Popularized" doesn't seem like quite the right word though, I don't disagree with the point, but if I shout "Hey everyone let's write regex's" at the office people throw stationary at me, which is not true of other popular things!
I took a look at Raku, which claims be a better Perl maybe, or closely related but more modern, it certainly looks nice. Although i am a big fan of typed languages, Raku piqued my interest.
my Int $a = 42; # ok
my Int $a = "foo"; # Type check failed in assignment to $a; expected Int but got Str ("foo")
(But that might not solve that problem? Maybe the problem is mostly about using same-character delimiters for strings.)
And I guess that’s why Perl is so flexible with regards to delimiters and such.
What you have gained is that the regex is now much easier to read.
That's like "in theory we need 4 bytes to represent Unicode, but in practice 3 bytes is fine" (glances at universally-maligned utf8mb3)
Only in multiline mode does it match EOL characters, but it does still not appear to consume them. In fact, I cannot construct a regex that captures the last character of one line, then consumes the newline, and then captures the first character of the next line, while using $. The capture group simply ends at $.
^ - takes you to start of line $ - takes you to end of line
In nearly two decades of using regex I think this might be the first time I've heard of $ being end of string. It's always been end of line for me.
These people are I think not intending to say a newline character is permitted at the end of an e-mail address.
(Of course people using 'grep' would have different expectations for obvious reasons)
I think this series of comments might be clearest: https://news.ycombinator.com/item?id=39764385
but yeah seems like a real misunderstanding from “start/end of string” people
String is usually end of line, but not if you use stuff like `N`, to manipulate multi-line strings
Per POSIX chapter 9[1]:
9.2 … "The use of regular expressions is generally associated with text processing. REs (BREs and EREs) operate on text strings; that is, zero or more characters followed by an end-of-string delimiter (typically NUL). Some utilities employing regular expressions limit the processing to lines; that is, zero or more characters followed by a <newline>."
and 9.3.8 … "A <dollar-sign> ( '$' ) shall be an anchor when used as the last character of an entire BRE. The implementation may treat a <dollar-sign> as an anchor when used as the last character of a subexpression. The <dollar-sign> shall anchor the expression (or optionally subexpression) to the end of the string being matched; the <dollar-sign> can be said to match the end-of-string following the last character."
combine to mean that $ may match the end of string OR the end of the line, and it's up to the utility (or mode) to define which. Most of the common utilities (grep, sed, awk, Python, etc) treat it as end of line by default, since they operate on lines by default.
THERE IS NO SINGLE UNIVERSAL REGULAR EXPRESSION SYNTAX. You cannot reliably read or write regular expressions without knowing which language & options are being used.
[1] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
His latest on the topic is cool too: https://www.youtube.com/watch?v=ys7yUyyQA-Y
He's has quite a lot of content that HN folks might be interested in I think, like the reality and woes of consulting[3]
[0] https://www.youtube.com/@RobertElderSoftware
[1] https://blog.robertelder.org/
Today most important bit of information is knowledge that implementations differ and I made a habit of pulling reference sheet for a thing I work with.
E.g. Emacs Regexp annoyingly doesn’t have word in form of “\w” but uses “\s_-“ (or something no reference sheet on screen) as character class (but Emacs has the best documentation and discoverability - a hill I’m willing to die on)
Some utilities require parenthesis escaping and some not. Sometimes this behavior is configurable and sometimes it’s not.
I lived through whole confusion, annoyance, denial phase and now I just accept it. Concept is the same everywhere but flavor changes.
My brain thinks in Perl's regex language and then I have to translate the inconsistent bits to the language I'm using. Especially in the shell - I'm way more likely to just drop a perl into the pipeline instead of trying to remember how sed/grep/awk (GNU or BSD?) prefer their regex.
https://pcre.org/current/doc/html/pcre2compat.html
https://en.wikipedia.org/wiki/Perl_Compatible_Regular_Expres...
https://stackoverflow.com/questions/70273084/regex-differenc...
> Perl is kind of designed to make awk and sed semi-obsolete.
Learning Perl today would be a very different experience. I don't think it would catch me as readily as it did back then. But it doesn't matter - it's embedded into me at a deep level because I learned it through a strong drive of fascination and infatuation.
As for the regex themselves? It's powerful and solved a lot of the problems I was trying to solve, was built fundamentally into Perl as a language, so learning it was just an easy iterative process. It didn't hurt that the particular period of time when I learned Perl/regex the community was really big on "leetcode" style exercises, they just happened to be focused around Perl Golf, being clever in how you wrote solutions to arbitrary problems, and abusive levels of regex to solve problems. We were all playing and play is a great way to learn.
Regex, useful in any job...
if you know where to find something no point in knowing it.
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
https://chat.openai.com/share/696f7046-7f43-4331-b12b-538566...
chatgpt-3.5:
^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$
https://chat.openai.com/share/aaa09ae8-3fd9-4df7-a417-948436...
I know an email has to have a domain name after the @ so I know where to send it.
I also know it has to have something before the @ so the domain’s email server knows how to handle it.
But do I care if the email server is supports sub addresses, characters outside of the commonly supported range (eg quotation marks and spaces), or even characters which aren’t part of the RFC? I do not.
If the user gives me that email, I’ll trust them. Worst case they won’t receive the verification email and will need to double check it. But it’s a lot better than those websites who try to tell me my email is invalid because their regex is too picky.
Though I guess adding a check for invalid dot patterns might be worthwhile.
[1] https://html.spec.whatwg.org/multipage/input.html#email-stat...
[2] https://datatracker.ietf.org/doc/draft-ietf-emailcore-as/
I'd be more emphatic that you shouldn't rely on regexes to validate emails and that this should only be used as an "in the form validation" first step to warn of user input error, but the gist is there
> This regex is *practical for most applications* (??), striking a balance between complexity and adherence to the standard. It allows for basic validation but does not fully enforce the specifications of RFC 5322, which are much more intricate and challenging to implement in a single regex pattern.
^ ("challenging"? Didn't I see that emails validation requires at least a grammar and not just a regex?)
> For example, it doesn't account for quoted strings (which can include spaces) in the local part, nor does it fully validate all possible TLDs. Implementing a regex that fully complies with the RFC specifications is impractical due to their complexity and the flexibility allowed in the specifications.
> For applications requiring strict compliance, it's often recommended to use a library or built-in function for email validation provided by the programming language or framework you're using, as these are more likely to handle the nuances and edge cases correctly. Additionally, the ultimate test of an email address's validity is sending a confirmation email to it.
'I'm writing a nodejs javascript application and I need a regex to validate emails in my server. Can you write a regex that will safely and efficiently match emails?'
GPT4 / Gemini Advanced / Claude 3 Sonnet
GPT4: `const emailRegex = /^[a-zA-Z0-9._-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$/;` Full answser: https://justpaste.it/cg4cl
Gemini Advanced: `const emailRegex = /^[a-zA-Z0-9.!#$%&'+/=?^_`{|}~-]+@[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?(?:\.[a-zA-Z0-9](?:[a-zA-Z0-9-]{0,61}[a-zA-Z0-9])?)$/;` Full answer: https://justpaste.it/589a5
Claude 3: `const emailRegex = /^([a-zA-Z0-9._%-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,})$/;` Full answer: https://justpaste.it/82r2v
So a Internet email address match pattern has to be: "..*@..*", anything else can reject otherwise valid addresses.
That however does not account for earlier source routed addresses, not the old style UUCP bang paths. However those can probably be ignored for newly generated email.
I regularly use an email address with a "+" in the host part. When I used qmail, I often used addresses like: "foo-a/b-bar-tat@DOMAIN". Mainly for auto filtering received messages from mailing lists.
https://devcraft.io/2021/05/04/exiftool-arbitrary-code-execu...
> "a\ > ""
> The second quote was not escaped because in the regex $tok =~ /(\\+)$/ the $ will match the end of a string, but also match before a newline at the end of a string, so the code thinks that the quote is being escaped when it’s escaping the newline.
Nonsense. And you know it.
First, you need to know what to find, before knowing where to find it. And knowing what to find requires intricate knowledge of the thing. Not intricate implementation details, but enough to point yourself in the right direction.
Secondly, you need to know why to find thing X and not thing Y. If anything, ChatGPT is even worse than google or stackoverflow in "solving the XY problem for you". XY is a problem you don't want solved, but instead to be told that you don't want to solve it.
Maybe some future LLM can also push back. Maybe some future LLM can guide you to the right answer for a problem. But at the current state: nope.
Related: regexes are almost never the best answer to any question. They are available and quick, so all considered, maybe "the best" for this case. But overall: nah.
You start with what you do know, asking leading questions and being clear about what you don't, and you build towards deeper and deeper terminology until you get to the point where there are docs to read (because you still can't trust them to get the specifics right).
I've done this on a number of projects with pretty astonishing results, building stuff that would otherwise be completely out of my wheelhouse.
I was talking about X-Y on a higher level though. Architecture, Design Patterns, that kind of stuff. LLMs are (still?) particularly bad at this. Which is rather obvious if you think of them as "just" statistical models: it'll just suggest what is done most often in your context, not what is current best for your context.
Here's the explanation for $ in the perlre docs:
$ Match the end of the string
(or before newline at the end of the
string; or before any newline if /m is
used)This is a serious misconception. Perl is far, far from dead. The constant activity of the gargantuan CPAN library more than demonstrates very much the opposite.
I would say Perl and its community has done quite well considering it hasn't had the same mountain of corporate funds thrust into it like more highlighted have. Mainstream ain't everything.
That's one of the benefits of a complete rethink/rewrite, you can learn from the fact that the old behavior surprised people.
It seems so obvious that's the opposite of what they should have defaulted to, that it clearly should have been ^ and $ for lines, and ^^ and $$ for the string, since like ((1)(2)(3)):
^^line1$\n^line2$\n^line3$\n$
[1]: That, and it's not anywhere, while Perl 5 is everywhere.
I consider sed to be the baseline. If you can do sed you can do anything but it’s seriously limited.
The regular expression engines available in most mainstream languages go well beyond what is specified in POSIX though. An interesting example is named capturing group in Python, e.g., (?P<token>f[o]+).
Oddly, there are no backreferences in POSIX EREs.
Quoting from <https://pubs.opengroup.org/onlinepubs/9699919799.2008edition...>:
> It was suggested that, in addition to interval expressions, back-references ( '\n' ) should also be added to EREs. This was rejected by the standard developers as likely to decrease consensus.
Updated my comment to present a better example that avoids back-references. Thanks!
Edit: oh, you mean via regex engines available in GNU tools; I am dumb. Hmm... is there no GNU extension with PCRE?
Maybe there's some other implementation of sed that supports PCREs but that would really be an extension of that implementation of sed rather than a property of sed.
And maybe there's some GNU tool that uses PCREs, but that GNU tool would not be GNU sed, so it would not be a relevant property.
Anyway, they probably should have said BREs or EREs rather than "sed"...
It's a problem for Linux that it can't move on. The cold dead hand of gnu is firmly around the community's neck.
Android, Busybox and Muscl have entered the chat
I do wonder though what's the highest number of different regex syntaxes I've ever encountered (perhaps written?) within a single line: bash, grep and sed are never not in a "hold my beer" mood!
I've got "hold my beer" commits in .net - I've balanced brackets. I believe that's impossible in sed and grep. If I were going to write a json parser in a script, then a) stop me and b) it's got to be in powershell.
They seem to improve. Negative lookbehind isn't missing anymore [1]. But still lack the handy \Q and \E to escape stuff [2].
Your comment is missing a trigger warning, lol. But seriously, this is one of my flags for "this should probably be a script, or an awk or perl one-liner."
I feel it's like driving a rental car. It behaves slightly different than your own car, some features missing, some other features added, but in general, most of the things are pretty similar.
A lot of systems implemented PCRE, including JavaScript, since Perl extended the POSIX system with many useful extensions. IIRC, re2 tries to reign in on some of the performance issues and quirks the original systems had, while implementing the whole thing in Go.
edit: Did not realize re2 predated go ...
re2 adds a legitimate option to the menu of using NDFAs, which have the disadvantage of not supporting backreferences, but have the advantage of having constrained complexity of scanning a string. This does not come for free; you can conceivably end up with a compiled regexp of very large size with an NDFA approach, but most of the time you won't. The result may be generally slower than a PCRE-type approach, but it can also end up safer because you can be confident that there isn't a pathological input string for a given regexp that will go exponential.
This is one of those cases where ~99% of the time, it doesn't really matter which you choose, but at the scale of the Entire Programming World, both options need to be available. I've got some security applications where I legitimately prefer the re2 implementation in Go because it is advantageous to be confident that the REs I write have no pathological cases in the arbitrary input they face. PCRE can be necessary in certain high-performance cases, as long as you can be sure you're not going to get that pathological input.
RE engines don't quite engender the same emotions as programming languages as a whole, but this is not cheerleading, this is a sober engineering assessment. I use both styles in my code. I've even got one unlucky exe I've been working with lately that has both, because it rather irreducibly has the requirements for both. Professionally annoying, but not actually a problem.
* Finite automata based regex engines don't necessarily have to be slower than backtracking engines like PCRE. Go's regexp is in practice slower in a lot of cases, but this is more a property of its implementation than its concept. See: https://github.com/BurntSushi/rebar?tab=readme-ov-file#summa... --- Given "sufficient" implementation effort (~several person years of development work), backtrackers and finite automata engines can both perform very well, with one beating the other in some cases but not in others. It depends.
* Fun fact is that if you're iterating over all matches in a haystack (e.g., Go's `FindAll` routines), then you're susceptible to O(m * n^2) search time. This applies to all regex engines that implement some kind of leftmost match priority. See https://github.com/BurntSushi/rebar?tab=readme-ov-file#quadr... for a more detailed elaboration on this point.
Good on you.
So, yes, at least someone (me) considers regex to be standardized in several published de jure standards.
[0] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2013/n3690.pdf#chapter.28
[1] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1_chap09.html
[2] https://262.ecma-international.org/14.0/#sec-regexp-regular-expression-objectsThe semantics of ^ and $ is based on lines - whether single-line or multi-line mode. For string based semantics - which you could also think of as entire file if you are dealing with files - use \A and \Z or their equivalents.
[1] Both interpretations have their merits. If you transmit text over a serial connection, it is useful to have a newline as line terminator so that you know when you received a complete line. If you put text into text files, it might arguably be easier to look at a newline as a line separator because then you can not have a invalid last line. On the other hand having line terminators in text files allows you to detect incompletely written lines.
https://homakov.blogspot.com/2012/05/saferweb-injects-in-var...
> A pattern is a sequence of pattern items. A caret '^' at the beginning of a pattern anchors the match at the beginning of the subject string. A '$' at the end of a pattern anchors the match at the end of the subject string. At other positions, '^' and '$' have no special meaning and represent themselves.
https://www.lua.org/manual/5.3/manual.html#6.4.1
Lua's pattern matching is much simpler than regexes though.
> Unlike several other scripting languages, Lua does not use POSIX regular expressions (regexp) for pattern matching. The main reason for this is size: A typical implementation of POSIX regexp takes more than 4,000 lines of code. This is bigger than all Lua standard libraries together. In comparison, the implementation of pattern matching in Lua has less than 500 lines.
There's an additional caveat: if you use the optional "init" parameter to specify an offset into the string to start matching, the ^ anchor will match at that offset, which may or may not be what you expect.
I find it hard to imagine any other expectation passing the rubber duck test.
"Oh, so you expected the match to always fail, no matter what the string was?"
Don't get me wrong, it's certainly far more useful as it is, I'm glad it works this way.
In raku (aka perl6) Regexes were reinvented by Larry Wall (the creator of perl which made perlRE the de facto regex standard)
Here's what he does with $:
(https://docs.raku.org/language/regexes#Start_of_string_and_e...)
* The $ anchor only matches at the end of the string
* The $$ anchor matches at the end of a logical line. That is, before a newline character, or at the end of the string when the last character is not a newline character.
>>> import re
>>> bool(re.search('abc$', 'abc'))
True
>>> bool(re.search('abc$', 'abc\n'))
True
>>> bool(re.search('abc$', 'abc\n\n'))
False
>>> bool(re.search('abc$', 'abc '))
False
>>> bool(re.search('abc$', 'abc\t'))
False
>>> bool(re.search('abc$', 'abc\r'))
False
>>> bool(re.search('abc$', 'abc\r\n'))
False unless person_id =~ /^\d+$/
abort "Bad person ID"
end
sql = "select * from people where person_id = #{person_id}"
In addition to injection attacks, this also can bite people when parsing headers, where a bad header is allowed to sneak past a filter. $ ruby -e 'x = "25" ; if x =~ /^\d+$/ ; puts "yes" ; else ; puts "no" ; end'
yes
$ ruby -e 'x = "25\n" ; if x =~ /^\d+$/ ; puts "yes" ; else ; puts "no" ; end'
yes
$ ruby -e 'x = "a25\n" ; if x =~ /^\d+$/ ; puts "yes" ; else ; puts "no" ; end'
no
Also, you'd want to use something that parameterizes the query with '?' (I use the Sequel gem) instead of just stuffing it into a sql string. ruby -e 'x = "a\n25\n" ; if x =~ /^\d+$/ ; puts "yes" ; else ; puts "no" ; end'
yes
Good to know.The second line should always be no, which if you use `\A\d+\z`, it will be.
$ ruby -e 'x = "25\n; delete from people" ; if x =~ /^\d+$/ ; puts "yes" ; else ; puts "no" ; end'
yeshttps://devcraft.io/2021/05/04/exiftool-arbitrary-code-execu...
https://www.postgresql.org/docs/current/functions-matching.h...
my $string = "cat\n";
/cat$/s -> true
/cat\Z/s -> true
/cat\z/s -> falseThe end of line character is usually the standard Windows \r\n.
Yes, that means if you want to really match the end of line you have to match "\r$". So broken.
I bet if someday VS Code's Windows build ships with LF default on new installations, people won't even notice.
I mean, at some point it did matter what the OS did when you pressed the "Enter" button. But this isn't really the case much anymore. VS Code catches that keypress, and inserts whatever "files.eol" is set to. Sublime does the same. I didn't check, but I assume every other IDE has this setting.
Similarly, the HTML spec, which is pretty nuts, makes browsers normalize my enters to LF characters as I type into this textarea here (I can check by reading the `value` property in devtools), but when it's submitted, it converts every LF to a CRLF because that's how HTML forms were once specced back in the day. Again though, what my OS considers to be "the standard newline" is simply not considered at all. Even CMD.EXE batch files support LF.
I don't really type newlines all that much outside IDEs and browsers (incl electron apps) and places like MS Word, all of which disregard what the OS does and insert their own thing. Maybe the terminal? I don't even know. I doubt it's very consequential.
EDIT: PSA the same holds for backslashes! Do Not Use Backslashes. Don't use "OS specific directory separator constants". It's not 1998, just type "/" - it just works.
I don't know if it is the case on Windows 11, but I have surely been bitten by CMD batch files using LF line endings. I don't remember the exact issue but it may have been the one bug affecting labels. [1]
[1]: https://www.dostips.com/forum/viewtopic.php?t=8988#p58888
As with '/', they really ought to do this some day but won't.
And if you believe \r\n is the way to go, please make sure \n\r also works as they should have the same results. (or \r\n\r\r\r\r for that matter)
\r - returned teletype head to the start of a line
\n - move paper one line down
> The sequence CR+LF was commonly used on many early computer systems that had adopted Teletype machines—typically a Teletype Model 33 ASR—as a console device, because this sequence was required to position those printers at the start of a new line. The separation of newline into two functions concealed the fact that the print head could not return from the far right to the beginning of the next line in time to print the next character. Any character printed after a CR would often print as a smudge in the middle of the page while the print head was still moving the carriage back to the first position. "The solution was to make the newline two characters: CR to move the carriage to column one, and LF to move the paper up."[2] In fact, it was often necessary to send extra padding characters—extraneous CRs or NULs—which are ignored but give the print head time to move to the left margin. Many early video displays also required multiple character times to scroll the display.
The handle does 2 things: return and feed. You can also just return by not pulling all the way or the other way around depending on the design
Line Feed (\n) - move down 1 line
Essentially, when you're sending commands to teletypes, that's how you'd have to do it.
(For example, most text-based internet protocols such as HTTP use \r\n as a line separator/terminator, so refusing to put the \r will be incompatible for no reason.)
The rationale was probably "it should be easier to match input strings" and now it's harder for everyone.
Brad Freidl's Mastering Regular Expressions is a good book to read if you want to stop being surprised/lost.
I'll admit I stopped at the dive into DFA/NFA engine details.
Has anyone confirmed this behaviour directly against the runtimes/languages? Newlines at the end of a string are certainly something that could get lost in transit inside an online service involving multiple runtimes.
I've also run that locally against "go1.22.1 darwin/arm64", "go1.21.5 windows/amd64", and "go1.21.0 linux/amd64" with the same result.
In what way could newlines at the end of a string "could get lost in transit"?
There are plenty of other ways, too; bugs happen.
> The ^ and $ language elements indicate the beginning and end of the input string. The end of the input string can be a trailing newline \n character.
Beyond what's in the OP, that includes RE2, Hyperscan, D's std.regex, ICU, Perl, Python's third party `regex` package, and `regress`.
Matching EOL feels natural for every line-based process.
What I find way more annoying is escaping characters and writing character groups. Why can't all regex engines support '\d' and '\w' and such? Why, in sed, is an unescaped '.' a regex-dot matching any character, but an unescaped '(' is just a regular bracket?
It is because sed predates the very influential second generation Extended Regular Expression engine and by default uses the first generation Basic Regular Expression engine. So really it is for backwards compatibility.
http://man.openbsd.org/re_format#BASIC_REGULAR_EXPRESSIONS
you can usually pass sed a -r flag to get it to use ERE's
Actually I don't really know if BRE's predate ERE's or not. I assume they do based on the name but I might be wrong.
The work originally came from work by Stephen Cole Kleene in the 1950s. It was introduced into Unix fame via the QED editor (which later became ed (and sed), then ex, then vi, then vim; all with differing authors) when Ken Thompson added regex when he ported QED to CTSS (an OS developed at MIT for the IBM 709, which was later used to develop Multics, and hence lead to Unix).
Also the "grep" command got its name from "ed"; "g" (the global ed command) "re" (regular expression), and "p" (the print ed command). Try it in vi/vim, :g/string/p it is the same thing as the grep command.
for portability, -E is the POSIX flag for the same thing
My reasoning would be if the file is transmitted and gets truncated nobody would know for sure if it does not end a new line. Brownie points if this is code end has a comment that the files ends there.
The article calls computer languages platforms but the are computer languages. Bash is not included. Weird. I believe the most common use of regular expressions is the use of grep or egrep with bash or some other shell but, who knows. Maybe I am hanging with the wrong crowd.
- The JS/Go/Rust family, which treats $ like \z and does not support \Z at all
- The Java, .NET, PHP, Python family, which treats $ like \Z and may or may not (Python) support \z.
\Z does away with \n before the end of the string, while \z treats \n as a regular character. For multiline $ the distinction doesn't matter, because \n is the end.
Really the only deviation from the rule is Python's \Z, which is indeed weird.
They go only on the tokenizer, if they go somewhere at all.
As long as sharing anecdata, in 30 years, it's almost the only way I use it.
It's incredible for slicing and dicing repetitious text into structure. You generally want some sort of Practical Extraction and Reporting Language, the core of which is something like a regular expression, generally able to handle the, well, irregularity.
Most recent example (I did this last week) was extracting Apple's app store purchases from an OCR of the purchase history available through Apple's Music app's Account page that lets you see all purchases across all digital offerings, but only as a long scrolling dialog box (reading that dialog's contents through accessibility hooks only retrieves the first few pages, unfortunately).
Each purchase contains one or more items and each item has one or more vertical lines, and if logos contain text they add arbitrary lines per logo.
A good match and sub match multi-line regex folds that mess back into a CSV. In this case, the regex for this was less than an 80 char line of code and worked in the find replace of Sublime Text which has multiline matching, subgroups, and back references.
Another way to do this is something like a state match/case machine, but why write a program when you can just write a regular expression?
(/m enables multiline mode)
It must be some kind of philosophical objection because there's no way something with as much water under the bridge as Python simply hasn't got around to it.
> So if you're trying to match a string without a newline at the end, you can't
only use $ in Python! My expectation was having multiline mode disabled
wouldn't have had this newline-matching behavior, but that isn't the case.
I would argue this is correct behavior, a "line" isn't a "line" if it doesn't end with \n.[1] > 3.206 Line - A sequence of zero or more non- <newline> characters plus a terminating <newline> character.
[1] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...$ man pcre2syntax
Where you'll find the following block under ANCHORS AND SIMPLE ASSERTIONS:
$ end of subject
also before newline at end of subject
also before internal newline in multiline mode
So all the cases of "newline at/before end of subject" are covered here. Then, the question becomes "what is a subject?" Is it line-by-line? Are newlines included? What if we want multiline matching? That's where re.MULTILINE comes from, it's not "multiline matching" (sort of) it's "what is the subject of the regular expression that we're matching against"The other results that differ across engines seem to be because people either don't understand regex or because the POSIX description of how to deal with such an input and config was ill-defined.
I'm not too surprised by PHP and Python getting it wrong. Java and C# is a slight surprise though.
No, you're not, except for this weird corner case where `$` can match before the last `\n` in a string. It's not just any `\n` that non-multiline `$` can match before. It's when it's the last `\n` in the string. See:
>>> re.search('cat$', 'cat\n')
<re.Match object; span=(0, 3), match='cat'>
>>> re.search('cat$', 'cat\n\n')
>>>
This is weird behavior. I assume this is why RE2 didn't copy this. And it's certainly why I followed RE2 with Rust's regex crate. Non-multiline `$` should only match at the end of the string. It should not be line-aware. In regex engines like Python where it has the behavior above, it is only "partially" line-aware, and only in the sense that it treats the last `\n` as special.One could certainly have a debate whether this behavior is too strongly tied to the origins of regular expressions and now does more harm than good, but I am not convinced that this would be an easy and obvious choice to have breaking change.
I think you've kind of missed the point. Sure if `$` in non-multiline mode means "end of line" the behaviour might be reasonable. But the big error is that people DO NOT EXPECT `$` to mean "end of line" in that case. They expect it to mean "end of string". That's clearly the least surprising and most useful behaviour.
The bug is not in how they have implemented "end of line" matching in non-multiline mode. It's that they did it at all.
> Both ^ and $ always match at start or end of lines
This is trivially not true, as I showed in my previous example. The haystack `cat\n\n` contains two lines and the regex `cat$` says it should match `cat` followed by the "end of a line" according to your definition. Yet it does not match `cat` followed by the end of a line in `cat\n\n`. And it does not do so in Python or in any other regex engine.
You're trying to square a circle here. It can't be done.
Can you make sense of, historically, why this choice of semantics was made? Sure. I bet you can. But I can still evaluate the choice on its own merits today. And I did when I made the regex crate.
> but I am not convinced that this would be an easy and obvious choice to have breaking change.
Rust's regex crate, Go's regexp package and RE2 all reject this whacky behavior. As the regex crate maintainer, I don't think I've ever seen anyone complain. Not once. This to me suggests that, at minimum, making `$` and `\z` equivalent in non-multiline mode is a reasonable choice. I would also argue it is the better and more sensible approach.
Whether other regex engines should have a breaking change or not to change the meaning of `$` is an entirely different question completely. That is neither here nor there. They absolutely will not be able to make such a change, for many good reasons.
Sure, it takes a string which might be a line or multiple or whatever. Does not change the fact that $ matches at the end of a line. If you want the end of the string, use \Z.
This is trivially not true, as I showed in my previous example. The haystack `cat\n\n` contains two lines and the regex `cat$` says it should match `cat` followed by the "end of a line" according to your definition.
In multi-line mode it matches, in single-line mode it does not because there is a newline between cat and the end of the line. A newline is only a terminating newline if it is the last character, the newline after cat is not a terminating newline. You need cat\n$ or cat\n\n to match.
This only makes sense if re.search accepted a line to search. It doesn't. It accepts an arbitrary string.
I don't think this conversation is going anywhere. Your description of the semantics seems inconsistent and incomprehensible to me.
> A newline is only a terminating newline if it is the last character, the newline after cat is not a terminating newline. You need cat\n$ or cat\n\n to match.
The first `\n` in `cat\n\n` is a terminating newline. There just happens to be one after it.
Like I said, your description makes sense if the input is meant to be interpreted as a single line. And in some contexts (like line oriented CLI tools), that can make sense. But that's not the case here. So your description makes no sense at all to me.
Which is fine because lines are a subset of strings. And whether you want your input treated as a line or a string is decided by your pattern, use ^ and $ and it will be treated as a line, use \A and \Z and it will be treated as a string.
The first `\n` in `cat\n\n` is a terminating newline. There just happens to be one after it.
Look at where this is coming from. You do line-based stuff, there is either no newline at all or there is exactly one newline at the end. You do file-based stuff, there are many newlines. In both cases the behavior of ^ and $ makes perfect sense.
Now you come along with cat\n\n which clearly falls into the file-based stuff category as it has more than one newline in it but you also insist that it is not multiple lines. If it is not multiple lines, then only the last character can be a newline, otherwise it would be multiple lines.
And I get it, yes, you can throw arbitrary strings at a regular expression, this line-based processing is not everything, but it explains why things behave the way they do. And that is also why people added \A and \Z. And I understand that ^ and $ are much nicer and much better known than \A and \Z. Maybe the best option would be to have a separate flag that makes them synonymous with \A and \Z and this could maybe even be the default.
Where is this semantic explained in the `re` module docs?
This is totally and completely made up as far as I can tell.
This also seems entirely consistent with my rebuttal:
Me: What you're saying makes sense if condition foo holds.
You: Condition foo holds.
This is uninteresting to me because I see no reason to believe that condition foo holds. Where condition foo is "the input to re.search is expected to be a single line." Or more precisely, apparently, "the input to re.search is expected to be a single line when either ^ or $ appear in the pattern." That is totally bonkers.
> but it explains why things behave the way they do
Firstly, I am not debating with you about the historical reasoning for this. Secondly, I am providing a commentary on the semantics themselves (they suck) and also on your explanation of them in today's context (it doesn't make sense). Thirdly, I am not making a prescriptive argument that established regex engines should change their behavior in any way.
If you're looking to explain why this semantic is the way it is, then I'd expect writing from the original implementors of it. Probably in Perl. I wouldn't at all be surprised if this was an "oops" or if it was implemented in a strictly-line-oriented context, and then someone else decided to keep it unthinkingly when they moved to a non-line-oriented context. From there, compatibility takes over as a reason for why it's with us today.
If you do not specify multi-line, bar$ matches a lines ending in bar, either foobar\n or foobar if the terminating newline has been removed or does not exist. If you specify multi-line, then it will also match at every bar\n within the string. So it either treats your input as a single line or as multiple lines. You can of course not specify multi-line and still pass in a string with additional newlines within the string, but then those newlines will be treated more or less as any other character, bar$ will not match bar\n\n. The exception is that dot will not match them except you set the single-line/dot-all flag, bar\n$ will match bar\n\n but bar.$ will not unless you specify the single-line/dot-all flag.
I would even agree with you that it seems a bit weird. If you have a proper line without additional newlines in the middle, then multi-line behaves exactly like not multi-line. Not multi-line only behaves differently if you confront it with multiple lines and I have no good idea how you would end up in a situation where you have multiple lines and want to treat them as one unit but still treat the entire thing as if it was a line.
I still have not seen anything from you that makes sense of the behavior that `cat$` does not match `cat\n\n`. Like, I realize you've tried to explain it. But your explanation does not make sense. That's because the behavior is strange.
The only actual way to explain the behavior of $ is what the `re` docs say: it either matches at the end of the string or just before a `\n` that appears at the end of the string. That's it.
cat$, the $ matches the end of the line, the second \n, cat is not directly before that. I guess you want the regex engine to first treat the input as a multi-line input, extract cat\n as the first line, and then have cat$ match successfully in that single line? What about cat$ and dog$ and cat\ndog\n.
Ignoring compatibility concerns, I would want the regex engine to behave the same way RE2, Go's regexp package and Rust's regex engine behave. I remember specifically considering Cox's decision ~10 years ago when writing the initial implementation of the regex crate. I thought Perl's (and Python's) behavior on this point was whacky then and I still think it's whacky now. So I followed RE2's semantics.
The OP is right to be surprised by this. And folks will continue to be surprised by it for eternity because it's an extremely subtle corner case that doesn't have a consistent story explaining its behavior. (I know you have proffered one, but I don't find it consistent in the context of a general purpose regex engine that searches arbitrary strings and not just lines.)
Of course, compatibility is a trump card here. I've acknowledged that. Changing this behavior now would be too hard. The best you could probably do is some kind of migration, where you provide the more "sensible" behavior behind an opt-in flag. And then maybe Python 4 enables it by default. But it's a lot of churn, and while people will continue to be confounded by this so long as the behavior exists, it probably isn't a Huge & Common Deal In Practice. So it may not be worth fixing. But if you're starting from scratch? Yes, please don't implement $ this way. It should match the end of the string when 'm' is disabled and the end of any line (including end of string and possibly being Unicode aware, depending on how much you care about that) when 'm' is enabled.
ed -> sed
ed -> grep
The line oriented mature makes sense.There is some sed multi-line capability if one uses the hold space, but it is much easier to just use awk.
The real issue is that no language nowadays "just" implements BRE or ERE since both specs are lacking in features.
Most languages instead implement some variant of Perl's regex instead (often called PCRE regex because of the C library that brought Perl's regex to C), which as far as I can tell isn't standardized, so you get these subtle differences between implementations.
A reproducible example would be nice. I don’t understand what it is he cannot do. `re.search('$', 'no new lines')` returns a match.
re.match('^bob$', 'bob\n')
I didn't want the trailing newline to be included.
re.match('^bob$', 'bobs') → no
Most people would expect 'bob\n' not to match, because I used '$' and it has an extra character at the end, just like 'bobs'. In Python it does match because '\n' is a special case.
This, in a nutshell, is the sort of problem which renders fallacious the notion that you can unit-test your way to correct software.
Here’s an interesting project for typed regular expressions: https://news.ycombinator.com/item?id=12292389
``` re.match(text).extract().rstrip(“\n”) ```
The new line is just empty but not the first line anymore.
3.195 Incomplete Line
A sequence of one or more non-<newline> characters at the end of the file.
3.206 Line
A sequence of zero or more non-<newline> characters plus a terminating <newline> character.
courtesy of [0]. See also [1] for rationale on "text file": Text File
[...] The definition of "text file" has caused controversy. The only difference between text and binary files is that text files have lines of less than {LINE_MAX} bytes, with no NUL characters, each terminated by a <newline>. The definition allows a file with a single <newline>, or a totally empty file, to be called a text file. If a file ends with an incomplete line it is not strictly a text file by this definition. [...]
[0] https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...[1] https://pubs.opengroup.org/onlinepubs/9699919799/xrat/V4_xbd...
- VSCODE
- Visual Studio
- Notepad
C/C++ compilers that process "incomplete lines"
All of them.
Will my users consider it a bug if my program doesn't process "incomplete lines": Yes. (Yes, I have had this bug logged
against code that I wrote).
Will I consider it a bug if your program doesn't process "incomplete lines": Yes.Note that all of those three apps come from the Windows world where CRLF was indeed a line separator, and a file ending with CRLF was considered to have an empty line at the end.
But yes, you absolutely should accommodate for incomplete lines.
"cat$" with multiline enabled
"cat$" with multiline disabled
"cat\z"
"cat\Z"no matches
PSA: Regex security is particular to each implementation flavor. Please know the nuances of a particular kind and be unambiguously precise.
$ does not mean end of string in Python.