Regular Expressions – Mastering Lookahead and Lookbehind
rexegg.com
rexegg.com
I find myself writing simpler ones and tying them together with app code just for sanity's sake.
Or just use PEG, parser combinators, or other more readable parsing abstractions
String pattern = "^https+" // match the protocol at the beginning
+ "([a-zA-Z])+" // match the machine name
+ ...
Honestly, I use regular expressions because, even in such format expanded with comments, I haven't seen anything more readable after you get used to regex operators. I guess the closest would be the alternative format in CL-PPCRE. For instance: CL-USER> (cl-ppcre:parse-string "\\b\\d{1,3}\\.\\d{1,3}\\.\\d{1,3}\\.\\d{1,3}\\b")
(:SEQUENCE :WORD-BOUNDARY (:GREEDY-REPETITION 1 3 :DIGIT-CLASS) #\.
(:GREEDY-REPETITION 1 3 :DIGIT-CLASS) #\.
(:GREEDY-REPETITION 1 3 :DIGIT-CLASS) #\.
(:GREEDY-REPETITION 1 3 :DIGIT-CLASS) :WORD-BOUNDARY)
But then, any such form can get mouthful: CL-USER> (cl-ppcre:parse-string "((\\b[0-9]+)?\\.)?\\b[0-9]+([eE][-+]?[0-9]+)?\\b")
(:SEQUENCE
(:GREEDY-REPETITION 0 1
(:REGISTER
(:SEQUENCE
(:GREEDY-REPETITION 0 1
(:REGISTER
(:SEQUENCE :WORD-BOUNDARY
(:GREEDY-REPETITION 1 NIL (:CHAR-CLASS (:RANGE #\0 #\9))))))
#\.)))
:WORD-BOUNDARY
(:GREEDY-REPETITION 1 NIL (:CHAR-CLASS (:RANGE #\0 #\9)))
(:GREEDY-REPETITION 0 1
(:REGISTER
(:SEQUENCE (:CHAR-CLASS #\e #\E)
(:GREEDY-REPETITION 0 1 (:CHAR-CLASS #\- #\+))
(:GREEDY-REPETITION 1 NIL (:CHAR-CLASS (:RANGE #\0 #\9))))))
:WORD-BOUNDARY)Alas, we don't get such luxuries as named groups or static typing... Woe is me.
I think the automata theory behind them is more important to know than proficiency with specific regular expression implementations.
If the biomed or embedded programmer deliberately does not make use of that functionality, he is inefficient.
I'm finding these comments hilarious. No true programmer!
Look, I used to write web application in the 1990's with vim on computers with video cards that didn't have X drivers for them. I'm well versed in regular expressions, having used maybe a dozen flavors of them over the last 20 years. Being snooty about how useful regular expression should be (in your opinion) to the work of every other programmer out there isn't going to change the experiences of those others. I maintain that for web development there are quite often uses, but for scientific software (which is what I do), embedded, non-web based CRUD/LoB, and many other applications - it's just not what it used to be.
EDIT: turns out I was mixing up the tone of your comment with that of bmn__ down below so I replied more belligerent than your comment warrented - no offense, I'm just going to leave it up regardless.
Since I may use 2-3 [non-web] languages at the same time, viable IDE options may go down to zero. I like how regex and other vim-specific features empower my typing enough to not use what constrains me in my toolset.
There are few people more frustrating to algorithms engineers and embedded engineers than people like you who think you are the only type of programmers out there or that the programming you do is somehow superior(despite largely relying on math you learned at a decent secondary school).
val inputIdTypWidthPat = new Regex("""(?si)@(\w+)\s+(\w+)\((\d+)\)""", "id", "typ", "width")
val inputIdTypeWidthCheck = RxInputMatchGroups(inputIdTypWidthPat,
List(InputMatchGroups("""@COUNTRY_CODE char(2),""",
List(MatchGroups("""@COUNTRY_CODE char(2)""",
List("COUNTRY_CODE", "char", "2"))))))
def getIdTypWidth(s: String): (Option[(String, String, Int)], Int) = {
val om: Option[Match] = inputIdTypWidthPat.findFirstMatchIn(s)
if (om.isDefined) {
if (om.get.groupCount == 3) {
(Some(om.get.group(1), om.get.group(2), om.get.group(3).toInt), om.get.end)
} else (None, 0)
} else (None, 0)
}You don't even have to be able to code to make use of regular expressions. You can use regular expressions when searching and replacing in editors (even slightly barebones editors like gedit or kate). You can transform input data from almost any format into any other format using nothing but your editor and a series of replace statements. (No computations though.)
I think they should teach regex in high school. Many people working in non-IT office jobs could benefit from knowing regex, and I think it's really quick to learn this. (Now if only Excel's/Word's search/replace supported regex...)
You might think I am just talking about Microsoft's quirky implementation but even in the Linux-sphere it isn't consistent see:
http://www.greenend.org.uk/rjk/tech/regexp.html
You take a complex format string which was design to use the fewest characters instead of with clarity in mind, you then have every major application and library diverge on basic support and spec for features, and then you have all of them hack on support for UNICODE in their own unique way.
Regular Expressions likely won't ever die, but I for one would happily switch to an alternative with better readability, UNICODE support from day zero, and fewer niche features to keep things uniform. I'm tired of re-learning RegEx only to have everything I've learned either be forgot or not work the second I app switch.
Learn those, or at least the main differences between them, and the vast majority of the regular expression engines in software you use will become more recognizable.
The two main differences between various engines are which characters are "literal" and which characters are "magic" (Vim's engine is particularly annoying here), and how to write the "convenience character classes" (like what the shorthand for "alphanumeric character class" is). But these are minor issues, once you've learned how to write a regex, these are trivial to look up.
Knowledge of regular expressions transfer from one engine to another just fine.
You've already described a feature which has different syntax in one of the primary regex dialects I use (Emacs).
Most notably, the choice operator can either be ordered like in PEGs (if the first branch matches, the other isn't evaluated) or pick the branch that produces the longest match, CFG-like.
I'd agree with you if everything weren't so easy to look up.
^(\((?1)?\))$If you didn't know, a regexp engine with capture groups is already stronger than what in formal languages theory is called "regular expressions".
https://stackoverflow.com/questions/1732348/regex-match-open...
Also, you don't wanna spoon feed students, they'll never learn to fish. You would indeed have to go as far as implementing a regex engine or implementing I don't know, a certain finite automate in regex. I'm kidding but all you could realistically achieve would be the usage of a catalogue like command-line-fu or stackexchange unless the whole thing fits in a broader cs syllabus/curriculum.
[1] lrovocative statement: sed and awk are breaking the "do one function and it well" idea of unix.
But yes, lots of people do seem to do the "I suck at regex". Even when I notice people do crazy long winded transformations by hand which I then do within seconds. Still doesn't seem enough motivation for them to learn them properly.
Though in my experience, outside some edge cases where the format never changes (like matching a domain name in a URL) a regex is hell to maintain when you come back 3 years later.
There is also always the fun of people trying (and failing) to use regex in emails.
If you can't resole the domain part, return an error.
The most important parts to get comfortable with are:
* Capture groups and alternation
* Character sets
* Anchors (start and end of line)
* Common escape sequences (digit, word, whitespace)
* Repeat (any, one or more, n-m)
* Common flags (global, multi-line, case-insensitive)
If you need more than that, it's time to start evaluating other tools IMO. A lookahead here and there is okay, but I'd avoid them if possible.
The best way I found to learn regex was to take a set of inputs that I wanted to match, (and a set that I didn't) and play around on https://regex101.com/ until I got a pattern that did what I wanted. You'll very quickly start to learn the above bulletpoints, and before long you'll be able to write patterns without any reference.
If you find that your regexes are getting too large or unwieldy or difficult to understand (despite knowing the above bulletpoints) then you probably need a parser or some other more suitable tool.
https://www.cheatography.com/davechild/cheat-sheets/regular-...
It's like solving a puzzle that ends up eliminating work (by doing said work) extremely concisely. What better kind of puzzle is there?
Regular expression is useful for sure. What is terrible is that every language, shell and platform has different "styles and implementation". On windows, cmd, powershell, C#, sql server, etc all have their own styles. It's similar enough and at the same time different enough to drive you insane. Throw in linux with their shells, vi(m), perl, etc all using their own variants.
But the biggest problem is that regex is prone to "set it and forget it" issue. It's something we use once in a while and forget. Was it brackets or parentheses or braces for defining character ranges? Does . or + signify one or more? And ever try deciphering someone's undocumented multiline regex? Fun times.
Omg, thank you, that is the insight I needed and now I completely get it. Internet +1 for the day.
grep -Po '(?<=...)pattern'
Which also cuts out and prints the relevant part of the line. This saves a trip through cut, awk or perl. (-P is PCRE and -o is print only matched characters, which the lookarounds aren't a part of.) $ echo 'foo=5, Bar=3; x1=83, y=120' | grep -oP '\b[a-z]+=\K\d+'
5
120
further reading: https://stackoverflow.com/questions/11640447/variable-length... grep -oP '(?<=...)pattern(?=...)'
The caveat however is that the look{ahead,behind} pattern has to be of fixed length.You've made my day.
> \A(?=\w{6,10}\z)(?=[^a-z][a-z])(?=(?:[^A-Z][A-Z]){3})(?=\D\d).
And then these guys wonder why people hate regexes? The "now you have 2 problems" quote fit perfectly for that case.
The regular formalism is pretty neat though. There are alternative syntaxes (e.g. multiline regexps in Python) that are better suited for complex matchers.
If you escape them it should work, I guess.
My preference in such cases is for multiple separated or longer REs (which can be at least split in the surrounding code) and each part named or heavily commented. Of course it's always worthwhile to consider non-RE solutions if the problem can be broken down enough.
EDIT: Fixed typo
I'm not sure how you could fix that without introducing completely new characters or color-coding parts of the expression though.
It's much better in languages with regex literals like Ruby and JavaScript.
import pegs
echo "xzxy" =~ peg"""
B <- A 'x' 'y' / C
A <- '' / 'x' 'z'
C <- C 'w' / 'v'
"""
Stack overflowNim needs to let go of its toy parsing algorithm.
example:
$ # all words except those starting with 'c' or 'C'
$ echo 'Car Bat cod12 Map foo_bar' | grep -ioP '\bc\w+(*SKIP)(*F)|\w+'
Bat
Map
foo_bar
for more details: https://www.rexegg.com/backtracking-control-verbs.html#skipf...I personally feel that control verbs are bad additions to the regexp, even though I do know that it is not a big addition to the regexp engine itself (e.g. naturally extended from posesssive quantifiers like `a++` or atomic groups `(?>foo)`). Most uses of such verbs can be expressed with combined parsers and simpler regexps, in the much simpler and maintainable way.
the `-o` option allows to output only matching portion, the regex is meant to extract all words other than those starting with 'c' or 'C'
here's hopefully better example
$ # do something with words not surround by quotes
$ echo 'I like "mango" and "guava"' | perl -pe 's/"[^"]+"(*SKIP)(*F)|\w+/\U$&/g'
I LIKE "mango" AND "guava" const {sequence, suffix} = composeRegexp;
const maybe = suffix("?");
const oneOrMore = suffix("+");
const urlMatcher = sequence(
/^/,
"http"
maybe("s"),
"://",
maybe("www."),
oneOrMore(/[^ ]/),
/$/
);Hey, did you commit already? Still typing?
a b foo c
t.a = !!message.a;
t.b = String(message.b);
t.foo = message.foo;
t.c = tonum(message.c);