Regular Expressions make me feel like a powerful wizard: that's not a good thing
shkspr.mobi
shkspr.mobi
mandatory_leading_letters = "^\w+"
optional_suffix = "([-+.']\w+)*"
domain = "\w+"
domain_suffix = "([-.]\w+)*"
tld = "\.\w+([-.]\w+)*$"
regex = "{mandatory_leading_letters}{optional_suffix}@{domain}{domain_suffix}{tld}"
Then you could understand it? Seems to me the trouble isn't with regex but with the decision to write a regex without trying to make it understandable. It is also possible to minify your onto one line in some languages or otherwise obfuscate it; should we then condemn it?
If the author had to fix a multi-line regex, I hope that in the process, they understood it enough to break it up into pieces so that the next time it would be more possible to debug.
Regex are pretty neat. They are often not the solution. I would not want to use lookaround for example. But also, they are really useful in a lot of cases, and I would hate to not use them because they're capable of being obfuscated.
I have been writing regexes like this for 20+ years. The "well then just don't write line noise...?" realization comes easily for those of us that spent a few years "maturing out of Perl". Because first you stop writing Perl line noise, and then you realize that regexes are just another part of that line noise (because regex are so common in perl code).
Regular expressions DEFINITELY ARE the right tool if the problem that you have is that you need to parse a regular grammar (or near-regular grammar)!!!
Like, literally, they are 100% the right tool. Theoretically. Practically. Everything. The. Right. Tool.
And they are also the absolutely, terribly, completely wrong tool if you need to do anything other than parse a (very nearly) regular grammar.
We can write the regex like this, including all whitespace:
(?x)
^\w+ # mandatory leading letters
( [-+.'] \w+ )* # optional suffix
@
\w+ # domain
( [-.] \w+ )* # domain suffix
( \.\w+ ( [-.] \w+ )* )* #tld
$
Also, btw, I hope no one is really using this regex. It's wrong; for example it appears to be deliberately designed to fail on IDNs.When reading the code, unless you're specifically debugging the regular expression, you don't really need to get into the details of validating what the regular expression is doing.
I only recently started using it exclusively to parse grammar, when the rules got complex so did my regexes, but when I divided my regexes into groups each on their own variable and I formed whole a regex by concatenating these variables, regex was much more concise for me.
[1] I think that sticking even moderately complicated regexes in a modern codebase is probably not a great plan most of the time (lots of modern languages have easier-to-use options for accomplishing complex parsing tasks!), but given that regex libraries typically ship as part of the standard library and are extraordinarily well documented as a language, they're not a bad tool to reach for in a lot of cases.
I have also heard of another one as well. I suspect it will be quite a popular project for awhile. I wouldn't care if I never had to write another regex again.
> the very existence of RegEx101.com ought to bring shame on our industry.
> ...
> You have a desire to build something hard to debug.
Regex101 is there to make it easy to debug regexes. It would be MUCH harder to debug a complex parsing function than a regex with Regex101.
> You don't trust compilers.
I don't even understand the idea here. You need a compiler for the regex. Why would someone that doesn't trust compilers use regexes?
I encountered, this week, a 42-line function that determines if a number is a valid string representation of a double. The function is wrong and has been wrong for a decade or more. The code is obtuse using unclear state variables (multiple boolean flags) to accept or reject the string at various points. It could have been a regex, and we could have seen at (almost) a glance what was intended. There are several other functions in the same file doing similar things so this ended up being something like 300+ lines of both wrong and difficult to understand code that could have been maybe 30 total lines of relatively easy to read and debug regexes.
Compare for example something trivial like
[a-f0-9]{32}
with its unrolled form public boolean hashTest(String path) {
int runLength = 0;
int minLength = 32;
if (path.length() <= minLength + 2)
return false;
for (int i = 0; i < path.length(); i++) {
int c = path.charAt(i);
if ((c >= '0' && c <= '9') || (c >= 'a' && c <= 'f')) {
runLength++;
}
else if (runLength >= minLength) {
return true;
}
else {
runLength = 0;
}
}
return runLength >= minLength;
}Use a parser combinator library and your parser is much easier to debug, because everything's compositional and you can unit-test it.
Why do we allow all this reductive "why do we make this so hard" drivel? Why is it that people's attitude isn't, "wow, something I don't know, I should learn what this is..." instead of "STUPID HULK SMASH!!! HULK BRAIN HURT!!!! OWWWW!!!!"
Like, I'm sorry the author sucks. He doesn't know regular expressions very well as parts of his post are just plain wrong, and his regular expression itself looks unnecessarily convoluted.
Regular expressions exist because string parsing code fucking SUCKS. I started programming in Visual Basic, where there are no regexes and string parsing was a verbose hell of substring processing, array indices, and equivalency checks. Terse regular expressions can replace dozens of lines of word vomit, and the basic library of symbols isn't that hard to memorize. It is WAY easier to keep a regular expression in your head than the alternative.
For every non-trivial regex I write, I comment it heavily. I just put it together using string concatenation, like:
var r = "/[...]" + // comment goes here
"..." + // another comment
"(" +
"..." // yet another comment
"|..." // and another
and so on. Works like a charm, and I make sure to indent on the parentheses as well.Also in the past when a regex has various parts repeated, I'll define regex "parts" as variables, and then use those.
You can make regexes as clear as you like.
# Delete (most) C comments.
$program =~ s {
/\* # Match the opening delimiter.
.*? # Match a minimal number of characters.
\*/ # Match the closing delimiter.
} []gsx;
And I always do that to remind myself what I am doing it for.I also will paste one or more commented example lines of above my regexes so it is clear what kind of input you are expected to be processing, it always ends up being a time saver when debugging or updating code.
Anyway, the regex in question is extremely easy to read (the clue is in the name regular) (that's a word, followed by any number of groups that are one of these characters followed by a word... etc - you get the point). I think that regex is probably one of the easier ways of expressing the idea.
Regexes aren't 'skimmable', but they are (in a lot of cases, like this one) readable (because they are, yknow, pretty regular), so reading it slowly should be possible for anyone who knows what the basic rules are, which there aren't many of.
The chapter of the book "Real World Haskell" builds a simple CSV parser:
https://book.realworldhaskell.org/read/using-parsec.html
It may be adaptable for regex and for static analysis as well.
It's insane to write code to text search when a regular expression could be reasonably used instead. This article is objectively bad advice and should not be followed.
https://www.gnu.org/software/emacs/manual/html_node/elisp/Rx...
1. How a high-readability but consistent, structured, sugar-free syntax for regular expression looks like - as any code you'll hand-roll to replace your regexes will be strictly inferior to feeding Rx expression to a regex engine.
2. Why you'll still want to use plain regular expressions anyway. Rx expression grow large very quickly, so for any non-trivial problem, you'll quickly reach the point past which it's less readable than the raw regexp, by virtue of sheer size. And, again, whatever your language, your replacement for a regular expression is unlikely to beat Rx.
Those two points together form an argument that regular expressions are often the right tool for the job, and while seemingly requiring more concentration up front, they'll be easier to work with due to lower demand on your working memory.
In most string encoding it is hard or unergonomic to safely embed regexes or string literals in other regexes.
In this sense a regex is quite similar to SQL, just used for simpler operations generally.
It is to be noted that while regexes are used to parse (mostly) regular languages, the language of regex expressions is fully context-free.
Personally I dislike manually embedding context-free languages inside a regular encoding (strings) already embedded inside another context free language that could have just added support for structured regexes in the first place.
That's not even the case. You can define and compose rx expressions. This allows you to build larger more complex rx expressions from smaller simpler ones in much the same way as you manage the complexity of a large program by building it up from subroutines.
> build larger more complex rx expressions from smaller simpler ones in much the same way as you manage the complexity of a large program by building it up from subroutines.
Yes, and with it comes a problem: factoring out legos and composing the solution out of them reduces complexity, but it sacrifices locality. There's no free lunch: some things get easier because you get to ignore irrelevant detail, other things get harder, because the relevant details are all over the place. Humans have a limited working memory, so too much composition, or factoring along the wrong dimension to the problem you're solving, destroys readability - the solution no longer fits in your head, and you keep chasing pointers, constantly evicting one piece of the puzzle from your head to fit another one.
In this context, terseness and locality become desirable qualities. They make you expend more cognitive effort up front, but save you the working memory overhead of extra abstractions that come with composable pieces, and eliminate the cost of pointer chasing, as the whole thing is literally in front of your eyes all the time.
This is, to my understanding, the actual reason math-heavy fields (including mathematics itself) stick to dense equations built of single-character names chosen from several alphabets - it ultimately saves time. A single line in a math paper may fully express a complex thought which, were you to rewrite it in "clean code" style, would take 10 pages and involve several extra layers of abstractions.
I'm not advocating we should rewrite everything in APL (though I'm also not convinced it wouldn't be better on the net) - abstraction and composition are the fundamental tools that let us deal with complexity. But they have a cost, and sometimes that cost is too high. Regular expressions are, in my experience, usually a case of that.
With that said, I think the most ergonomic string searching tool I've used is the PEG implementation in Janet:
Also few off by ones and edge condition bugs lurking in “whatever”.
Where it makes sense to me is for instance ISS rewrite rules.
Just my 2 cents.
The problem is that they are often not use for simple computations, and therefore don't appear elegant.
They are, quite literally, like a flowchart. If it gets to be too big, then it gets difficult to understand without adding context to the parts (as some sibling comments have mentioned, by labeling the parts or by adding comments outside of the string definition).
Understanding Regex (and, by extension, DFAs) is very helpful, depending on your job. I used to teach Theory of Computing, and students would say "I'll never use this in the real world!", and yet only yesterday I was putting these very skills to use in my industry job in validating contract stipulations of a system.
I don't think it makes me feel like a powerful wizard... rather, I think I'm getting to spend time admiring a beautiful art exhibit!
[1] https://regexper.com/#%5E%5Cw%2B%28%5B-%2B.'%5D%5Cw%2B%29*%4...
The best I could do was to break the regexes down across multiple lines and comment each line, and break the handlers down into a hierarchy of smaller functions. Each function had a number of sample strings right in the code, and each file of code for handling the variations on output of a single command also had a unit test right in the same file which fed all the samples to the regexes. So the handling of the 40 different commands could be unit-tested very quickly with a single pytest command line. This was a godsend to help keep my change things-test things-change things again "loop" very quick.
The alternative might have been to actually write parsers, but I don't think that would have been more readable and it certainly wouldn't have been as quick.
I don't think language designers should get rid of regexes and I don't think programmers should stop using them. I do think there is definitely room for languages that implement regexes with more syntax. People might argue that this is just needless syntactic sugar, but so is spelling out keywords like "else." We ought to be looking at prior art. The history of programming languages is a very deep trashpile containing a lot of gold that can be mined with a little effort. SNOBOL's syntax for pattern matching was maybe too wordy by modern standards but I'm sure a lot of younger programmers would be pretty surprised to find how early a lot of things like regexes were invented, and how inspirational it can be to look at old languages.
Mathematical Naïve Set Theory is fundamentally flawed, otherwise Russell's Paradox wouldn't exist.
The thing is, with regex, maybe worst is better. It's the best we've got (for now), and I can't imagine parsing complex text without it. Manually parsing it (tokenize, etc) without regex is much harder.
The reasons: Writing parsing code by hand tends to be verbose, error prone (such as off by one errors), ad hoc, and full of temporary strings. Editor syntax highlighting (and also sites like regex101). Lastly regexes are matched in one go (match /no match) which is good for readability.
I do avoid making regexes long if I can.
The dirty secret of CS school is that your theory course is mostly about parsing until you get to turing machines (I joke, but only kind of)
[1] https://en.wikipedia.org/wiki/Parsing_expression_grammar
I find PEGs useful for actual CFG parsing, but tbh can’t see why I would avoid regex in favor of it in this case.
> I genuinely - and possibly misguidedly - believe that even something like
^\w+([-+.']\w+)*@\w+([-.]\w+)*\.\w+([-.]\w+)*$
> might just as well be written in BrainFuck.all those \w & \W's? c'mon, that's the easiest RE parsing task you could possibly confront!
it is really difficult to look at RE expressions, but it's also really hard to think of a better way
Jamie Zawinski (1997)
Dev who solves problem with regex, now has two problems.