Regex for Noobs – An Illustrated Guide
janmeppe.com
janmeppe.com
A few months ago I needed to add a regex for validating IPv4 addresses that were in a non-standard format. I had to do a little research, then I wrote & tested it, and promptly forgot 90% of what I had just learned.
It doesn't have to be this way. If most programming languages were designed like regular expressions they would be unusable. I can spend a year w/o writing a single line of Pascal or C++ and then it would come back immediately. I occasionally found regular expressions useful in my work, but only a couple of times could justify spending time on learning a little bit of it only to forget the next day. In a few other cases I simply googled a solution and used w/o understanding it. Most often I just write code instead, it is verbose but at least it's clear what it does.
I systemised as much as I could, but there was a lot of write practice when it came to learning trigonomic idendities and transidnetal functions. And not just the normal ones; I mean all of them ... both the inverserse and hyperbolic identities.
I did all of the problems and got extra ones.
In short, I knew the material backwards and forwards; I finished hour long exams in 20 minutes, and was usually the first or second person to leave.
What I am trying to say is that REGEX is similar to calculas and in my view requires significant focus and practice to master. Or you can muddle your way through it when you have to.
(also I love comments that point out amazing learning resources with full conviction)
I highly recommend the lessons. If you do them, I also highly recommend donating the recommended $4 via Paypal (or more), although it's technically free to use. It's well worth it.
Regexes get a bad rap because programmers who write otherwise maintainable code throw code hygiene out the window when writing long regexes. Your host language is much more complicated than the regex language, but writing shitty unreadable regex is, for whatever reason, acceptable. Some unfortunate and unnecessary limitations of typical regex libraries make the problem worse.
1. Regexes are typically written as one-line strings instead of as structu/red code. When writing C/Java, programmers understand that they should put each element of a sequence on a separate line and use indentation to visually signal branching and loops. But for some reason when writing character-processing programs that also have sequences, branching (|), and loops, programmers almost always golf it and put the whole complex program all on one line.
If you're writing a regex longer than 5 or 10 atoms, place each sequence on a separate line and use newlines+indentation to visually offset disjunctive choices (|) and loops (star).
Rule of thumb: you should be able to roughly sketch the rough shape of the finite automata by crossing your eyes and eyeballing the shape of the regex, just like you should be able to roughly sketch the control flow of a program by crossing your eyes and eyeballing the shape of the code.
2. Character group names are too terse (e.g., \s instead of "\whitespace" and \d instead of \digit), probably because of #1. Also, very few people give names for long subexpressions (e.g., factoring code out into functions with well-chosen names).
3. Gotos/try...catch (i.e., backtracking) are not used judiciously and aren't well-documented/tested. I often see backtracking or confusing mixes of lazy and greedy matching instead of just writing out a slightly longer disjunction.
4. Regexes are often used for languages that aren't even almost regular. A bit of backtracking is OK (gotos are sometimes OK), but if there's a lot of backtracking then you need to use a different class of languages/machines.
5. There's no way to embed non-regular matchers into a regex.
Due to the combination of 1-5, matching an email address with an optional recipient name (so something like "asdf@asdf.com" or something like "John Smith <asdf@asdf.com>") you need an insane regex that many people implement with lots of backtracking and so on all stuffed into a single line.
But something like this would work just fine and is much more readable:
\emailAddress := ...
# todo: need to support dashes in names.
\name := (
[a-zA-Z]*
\whitespace?
[a-zA-Z]*
)
# matches asdf@asdf.com or John Smith <asdf@asdf.com>
(
\emailAddress
)
OR
(
\name
\whitespace?
<
\emailAddress
>
)
where e.g., the implementation of \emailAddress could be written as a stand-alone parser in the host language. But even without digging into email address you can already see how this is way more readable than: \emailAddress|[a-zA-Z]*\s?[a-zA-Z]*\s?<\emailAddress>
Writing readable regular expressions shouldn't be difficult -- just treat the regex like any other piece of code and allow inter-op with the host language. But few people/libraries put in the effort, and for whatever reason golfing regexes in production code is considered acceptable even in orgs where you'd be fired for code golfing in the host programming language. lower = lpeg.R("az")
upper = lpeg.R("AZ")
letter = lower + upper
Personally I prefer though the terseness of regular expressions.> Personally I prefer though the terseness of regular expressions.
I think there are legitimate use-cases for both.
If you're quickly hacking out a small good-enough parser for something regular or "almost regular", terseness can be great.
However, if you're parsing a large regular language, terseness isn't really a benefit. Perl is on its bed for a reason, and overly terse regular expressions should die for a similar reason. Overly-clever write-only coding culture sucks.
But the terseness of regular expressions is basically terrible beyond maybe a few hundred characters. E.g., the following has no place in production code -- you might as well just include a binary:
(?:(?:\r\n)?[ \t])*(?:(?:(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t]
)+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:
\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(
?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[
\t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\0
31]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\
](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+
(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:
(?:\r\n)?[ \t])*))*|(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z
|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)
?[ \t])*)*\<(?:(?:\r\n)?[ \t])*(?:@(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\
r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[
\t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)
?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t]
)*))*(?:,@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[
\t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*
)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t]
)+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*)
*:(?:(?:\r\n)?[ \t])*)?(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+
|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r
\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:
\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t
]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031
]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](
?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?
:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?
:\r\n)?[ \t])*))*\>(?:(?:\r\n)?[ \t])*)|(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?
:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?
[ \t]))*"(?:(?:\r\n)?[ \t])*)*:(?:(?:\r\n)?[ \t])*(?:(?:(?:[^()<>@,;:\\".\[\]
\000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|
\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>
@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"
(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t]
)*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?
:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[
\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*|(?:[^()<>@,;:\\".\[\] \000-
\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(
?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)*\<(?:(?:\r\n)?[ \t])*(?:@(?:[^()<>@,;
:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([
^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\"
.\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\
]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*(?:,@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\
[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\
r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\]
\000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]
|\\.)*\](?:(?:\r\n)?[ \t])*))*)*:(?:(?:\r\n)?[ \t])*)?(?:[^()<>@,;:\\".\[\] \0
00-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\
.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,
;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|"(?
:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*))*@(?:(?:\r\n)?[ \t])*
(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".
\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t])*(?:[
^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\]
]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*\>(?:(?:\r\n)?[ \t])*)(?:,\s*(
?:(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(
?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[
\["()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t
])*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t
])+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?
:\.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|
\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*|(?:
[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".\[\
]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)*\<(?:(?:\r\n)
?[ \t])*(?:@(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["
()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)
?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>
@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*(?:,@(?:(?:\r\n)?[
\t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,
;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\.(?:(?:\r\n)?[ \t]
)*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\
".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*)*:(?:(?:\r\n)?[ \t])*)?
(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\["()<>@,;:\\".
\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])*)(?:\.(?:(?:
\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z|(?=[\[
"()<>@,;:\\".\[\]]))|"(?:[^\"\r\\]|\\.|(?:(?:\r\n)?[ \t]))*"(?:(?:\r\n)?[ \t])
*))*@(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])
+|\Z|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*)(?:\
.(?:(?:\r\n)?[ \t])*(?:[^()<>@,;:\\".\[\] \000-\031]+(?:(?:(?:\r\n)?[ \t])+|\Z
|(?=[\["()<>@,;:\\".\[\]]))|\[([^\[\]\r\\]|\\.)*\](?:(?:\r\n)?[ \t])*))*\>(?:(
?:\r\n)?[ \t])*))*)?;\s*)I never really got the point of regexes until I discovered capture groups. Pretty much everything I use them for is based on some problem like:
"Change date format from 'DD-MM-YY' to 'YYYY-MM-DD'"
"Keep only lines in the log file with a date in them"
"Swap the first name and last name of every contact in this list which has a comma in it"
"Rename all the files in this folder to strip out all the gubbins and put the useful stuff in a standard format"
And so on. You need a problem where doing it any other way than regex would be lengthy and painful and the regex itself is light and breezy. You build up little snippets or "phrases" that you can re-use easily to solve new problems.
Spaced repetition works!
Forcing myself to get tough regex working is what really made me grasp some of the more advanced concepts that unlock regex. It’s mainly just that we never need regex, and if you use it for a simple task you’re a jerk for introducing unnecessary complexity. Why would people learn it?
Also start using them to help you find and replace stuff in a text editor. They are very handy for converting for example a csv into code.
If you're in a hurry and just need to do some one-off slicing and dicing, there's not much more powerful than being able to say, eg:
$data = "foobarfoo";
$data =~ s/bar/baz/g;
print "$data\n";
versus python's more dramatic and formal: import re
data = "foobarfoo"
data = re.sub('bar', 'baz', data)
print data
For some reason, the Perl =~ notation sticks in my head
much more easily than the Python method.Of course you can string a bunch of sed commands together in a shell script as well, but I find that becomes unwieldy pretty quickly in a lot of situations.
In other languages a part of me always feels like if I'm using regex I'm doing something wrong, because I used it so liberally in Perl. It feels like a Perl-ism so I try to make a conscious effort not to carry those over as they tend to be the source of inefficiencies.
str.replace = replace(...)
S.replace(old, new[, count]) -> str
Return a copy of S with all occurrences of substring
old replaced by new. If the optional argument count is
given, only the first count occurrences are replaced.
(though the documentation turns out more misleading than re.sub's: this states it replaces all occurrences but really replaces all non-overlapping occurrences, which re.sub actually spells out properly).I found only one very weird edge case worth noting, in Perl the regex "s/ a* / x /g" (spaces added to prevent formatting) will turn the string "bac" into "xbxxcx" because the "a" technically means zero or more so it matches the spaces between the characters, not so in python it creates the string "xbxcx" because it matched the a in between one time and didn't count the empty spaces. Slightly less accurate results since * does mean zero or more so the space between counts as zero characters.
And coming to Perl vs Python regex differences, there are too many to count, Python's 3rd party module 'regex' is more comparable to Perl regex. For example, Python doesn't support possesive quantifiers, subexpression calls, \K, \G and so on
The gap between perl and python seems cleaner than the gap between python and anything-statically-typed, at least.
$data = "foobarfoo"
say $data =~ s/bar/baz/gr
say is like print, but adds a newline as well.The /r option does a non destructive replace and returns the result.
It's such a good resource for understanding and writing both simple and complicated regex.
As a "next level" site, also give https://rexegg.com a look.
but, I always try to add a warning when recommending - should use them only for the flavors supported (PCRE, Python, etc) I've seen many using it for cli tools and wonder why things like non-greedy or lookarounds don't work.
.NET's System.Text.RegularExpressions.Regex is one implementation that has no restrictions on lookbehind. Having used PowerShell for so long it now happens sometimes that I forget when writing regexes for other implementations, as it's really convenient at times.
a few do support it (for ex: Python's 3rd party regex module) and sometimes you could workaround with \K (similar to \zs in vim) [1]
And there are other frustrating differences between implementations, for example \g definition is very different between Ruby and Python, character set operations are not found everywhere, etc. Plus, BRE/ERE versions found in command line tools do not even support features like non-greedy and lookaround
[1] https://stackoverflow.com/questions/11640447/variable-length...
>For most people without a formal CS education, regular expressions (regex) can come off as something that only the most hardcore unix programmers would dare to touch.
my experience wasn't as daunting as this at all.
this is my story: basically every regex I ever wrote always worked the first time and I found it super easy. my intro was the Perl 5 "camel" book, i.e. Programming Perl. never heard of regexes before that.
If you find this current tutorial we're discussing tricky or daunting, maybe give the resource I just mentioned ("Programming Perl" for Perl 5) a go because for whatever reason the explanation was super simple and writing a regex was one of the easiest things I ever did no matter what I wanted to do. I didn't know about the concept itself before I read that book. I've used it to parse loads and loads of things, use it in my editors, etc. It's always so easy.
I don't know if you can find the "Programming Perl" book, I tried to look and found something like that here: https://www.cs.ait.ac.th/~on/O/oreilly/perl/prog/ch01_07.htm
you can maybe ignore the stuff about Perl and still sort of follow the tutorial.
just an alternative in case someone had their curiosity piqued by this illustrated guide, but it doesn't quite "click" for them,, and wanted to see the same thing written out in another teaching way. And a note that for me it was always not daunting at all!
No, I didn't. (at the time, bought the Perl book in the U.S., where I read it.)
I also found the various perldoc pages on regex to be very helpful resource that I referred to frequently back when I was writing perl. Notable pages are perlretut [0], which serves as an intro to regex, perlre [1] which goes into considerable depth, and perlreref [2] which is a handy quick reference for day-to-day work.
[0] https://perldoc.perl.org/5.30.0/perlretut.html
I think of regular expressions as a mini-programming language, and use loose fitting analogies
Anchors --> adding if conditions
Alternation --> Conditional OR
a(b|c)d = abd|acd --> a(b+c)d = abd+acd
Quantifiers --> string repetition and range operators
Dot metacharacter + quantifiers + alternation --> Conditional AND
Capture group and backreference --> variable
Subexpression call --> function
Character class --> sets
Lookarounds --> custom conditionals
Flags --> like command line options
So far, I've written books for Ruby and Python regular expressions, and BRE/ERE/PCRE/Rust for "GNU grep and ripgrep". Currently writing a book for GNU sed.The gory details of how to just match block comments are best outlined by this blog post detailing one kindred spirit's journey through the wilderness: https://blog.ostermiller.org/find-comment, although this wasn't particularly helpful to my situation, it was reassuring to know that I was not alone and might prove interesting to the crowd here.
I was introduced to them that way and it seemed very intuitive to me. I’m not sure if I just have a head for regular expressions, or if the way they were presented was especially good.
1) Regular expressions and "regex" are different things. Regex is a superset of regular expressions. But if you can understand regular expressions you can understand a huge number of regexes and have the tools to understand the rest of the syntax that the system uses.
2) There are many many many regular expression and regex dialects and engines. It's useful to learn the dialect you are trying to use.
3) Originally, regular expressions were used to describe a generator of strings. It can be helpful to approach writing regexes from this standpoint. It's not about what it matches, but what strings the regex can produce. I've had many students go from bewildered incomprehension to immediate eureka when I describe this.
4) There are only three rules that the beginning regex user has to know:
i) Concatenation - you can make a bigger generator by combining one generator after the other. For example: the regex 'a' and the regex 'b' can be concatenated into 'ab' and produces only the string 'ab'.
ii) Alternation - you can specify one character or another using the pipe character in most dialects '|'. For example the regex 'a|b' produces either the string 'a' or 'b'.
iii) Repetition - the character ' * ' is a postfix operator that means that the character that preceeds it can be repeated zero to an infinite number of times. For example: 'a* ' produces the empty string, 'a', 'aa', 'aaa' and so on.
5) By combining these (i.e. using concatenation) and the grouping operators '(' ')' you can write a regex for almost anything.
6) Nearly all of the other character you see in regexes that don't make sense are simply syntactic sugar that make it easier and shorter to write common expressions in a more compact form. Examples:
Writing lots of alternations of single characters (called character classes) - 'a|b|c|d|e|f' can be written as [a-f] - an expression that is the alternation of every character except the newline character can be written as '.' There's a whole host of these kinds of things.
Writing lots of different kinds of repetitions introduces lots of new operators - 'aa* ' can be written as 'a+' where '+' means 'repeat the previous character 1 or more times - 'a?' means 'repeat the previous character 0 or one times' - 'a{1,5}' means 'repeat the previous character 1 to 5 times' There's also an entire table of these things.
After that it's mostly practice and referring to the description of the operators for the dialect you are using.
I usually cover capture groupings (e.g. \1, \2, \3 or $1, $2, $3) later on if necessary as it requires people to have a good handle on the particulars of their languages way of handling those things. At this time I'll also over the '(?:.)' non-capture group operator if their dialect allows.
About as frequently I'll have to introduce '^' and '$' and begin and end string operators and also deconflict it from '[^.]'.
All the crazy-ass back/forward reference stuff I usually leave out as people almost never need or come across those except in pretty rare instances. In those cases they usually have enough regex under their belt that it's just another topic to pick up.
This doesn't get anybody to pro-wizard level and there's always edge cases and weird Perl-golf stuff that's out there, but it'll get most people to generally functional inside of a day or two and I'll just make myself available to spot answer questions or remind them where they forgot something.
> [0-9][0-9]* (what this pattern matches is left as an exercise to the reader)
Some more real world examples to bring all the concepts home would have been great.
The whole pattern therefore means one digit 0-9 followed by zero or many digits in the range 0-9.
so "3" would match "21" would match "123456" would match "" would not match
The obviously "correct" pattern for this behavior would be [0-9]+ The + means 1 or more and thus makes repeating the set obsolete.
Thanks!
[0|1|2|3|4|5|6|7|8|9][0|1|2|3|4|5|6|7|8|9]{0,} and [0-9][0-9]* and [0-9]+ and \d+ all do the same so they are not wrong as in the result will be correct.
>Regex For Noobs (like me!) - an illustrated guide
So I am painfully aware that I'm a noob too, I even call myself one! I wrote this little piece because I felt that regex can be kind of intimidating for a lot of people but that it's actually pretty OK and fun once you get the hang of it. To get people over that initial barrier is what I aim to do with the "guide".
I can recommend the articles by Julia Evans[1], which are pretty much like this. She works on a problem that is new to her and describes how she solved it or how she learned something new.
[1]: https://jvns.ca/
https://www.gnu.org/software/sed/manual/html_node/BRE-syntax...
You may wanna have a look at the PCRE flavor if you understand this you can work with all others (you'll only gonna miss some features)
The + is really just a shortcut for {1,} means 1 or more greedy the * is a shortcut for {0,} means 0 or more greedy Also while I'm at it the "even more correct" pattern ofc would be \d+ (at least in PCRE) \d stand for digit it shortcut for [0-9] Personally I would never write [0-9] I only ever use sets for partial ranges i.e. [5-9] or [a-f] Whenever you need the "full" set there is usually a short way to write it. This makes pattern shorter and easier to read. Another nice trick is to uppercase the shortcuts to invert it's meaning so \D would match any non-digit.
I recommend everybody who's interested in "looking behind the curtains" of regexp, their powers and limitations, to dive in a crash course of formal languages (Chomsky hierarchy). It can be useful to gain a feeling what a regexp can do and what it cannot.
https://github.com/ashchan/cheatsheets/blob/master/misc/regu...
For more advanced expressions I just do a quick google search for existing solutions before attempting to recreate the wheel.
Note: this is a new account with my real name, I used to have an account here already.
1. Don't put it all on one line
2. Indent nested constructs
3. Write comments
See for example Python's verbose regex syntax: https://docs.python.org/3/library/re.html#re.VERBOSE
But still, if you have to comment RE to improve readability is that really an improvement or one more thing to maintain?
My preference is to avoid RE in larger code bases.
See also: https://blog.codinghorror.com/regular-expressions-now-you-ha...