The Greatest Regex Trick Ever (2014)
rexegg.com
rexegg.com
I have a process table and I want to grep it for the phrase "banana":
ps auxww | grep banana
root 87 Jun21 0:26.78 /System/Library/CoreServices/FruitProcessor --core=banana
mikec 456 450PM 0:00.00 grep banana
Argh! It also greps for the grep for banana! Annoying!
Well, I'm sure there's pgrep or some clever thing, but my coworker showed me this and it took me a few minutes to realize how it works:
ps auxww | grep [b]anana
root 87 Jun21 0:26.78 /System/Library/CoreServices/FruitProcessor --core=banana
Doc Brown spoke to me: "You're just not thinking fourth dimensionally!" Like Marty, I have a real problem with that. But don't you see: [b]anana matches banana but it doesn't match 'grep [b]anana' as a raw string. And so I get only the process I wanted!
| grep -v grep
like in ps auxww | grep banana | grep -v grep(Also @thewakalix made a good suggestion to reverse the greps.)
the reversing the greps will be my new default behavior
$ printf 'banana\ngrep banana\n' | grep banana | grep -v grep
banana
$ echo $?
0
$ printf 'grep banana\n' | grep banana | grep -v grep
$ echo $?
1
To clarify my previous comment, `grep -v grep` exits with 0 if there was a line without "grep" in the output of `grep banana`.ps auxwww | grep banana | grep -v grep && echo ${PIPESTATUS[1]}
type of thing
Never thought of that. Nice.
So for example if you have a script that uses `pgrep -f banana` to search for a "banana" process, and you run that script twice in parallel, pgrep might see the other pgrep process and think "banana" is running even though it isn't.
I was bitten by this :)
$ echo [b]anana
[b]anana
$ touch banana
$ echo [b]anana
banana
You can escape the bracket and it will work: $ echo \[b]anana
[b]ananaEscaping works under zsh. My preferred method is single quotes:
echo '[b]anana'But nice regex though
Edit : someone already posted that solution https://news.ycombinator.com/item?id=27777901
ps auxww | grep b\\anana ps auxww|sed -n '/ doesnotexist /d;/banana/p'
When/if grep is not availableEmedded systems is one answer.
Toolchains for compiling systems from source is another answer, e.g., NetBSD's toolchain has sed, but not grep.
A third answer is install media. For example, NetBSD install kernels have ramdisks with sed but not grep.
A fourth answer is personal, customised systems. I create small systems that run from RAM. I run these on small computers with limited resources. When one of these computers first boots up, it may not have a "full" set of userland programs. It may not have grep. I am not inclined to use the limited space available to include grep at such an early stage if I can get by with sed.
Hope this answers your question.
It does, thanks for the detailed response.
edit: saw someone else posted this as well. should have known
One other solution would have been to run the regex twice, once to pick up all instances of Tarzan, and a second on the results of the first to filter out all instances of "Tarzan".
Some rare people can figure out:
\d{1,2}[-/]\d{1,2}[-/](\d{4}|\d{2})
but a dummy can figure out this:
(?<month> \d{1,2} ) [-/] (?<day> \d{1,2} ) [-/] (?<year> \d{4} | \d{2} )
Don't do too much in one operation whether it's regexes, SQL queries or OOP classes!
It’s still in a single batch of SQL (stored procedure in our case, so no additional network roundtrips), but the code is vastly clearer to read/maintain this way.
While maintaining/changing the SQL, comment in/out select-statements-as-printf-debugging, and comment in/out actual execution of the statements themselves.
These cursors would often contain [identifying object reference], [category of statement], [text of SQL statement to execute]. You would write a select statement to populate the cursor, then loop over the cursor to run all the statements in the order you wanted (drops, then user/role creates, then grants, or whatever the situation called for).
It's not about logical clarity, but practical maintainability given the (overall weak) state of tooling for database queries. Is it a bastardization of SQL to do something that "should be" done in another scripting language? Maybe, but there's a lot of power in giving the DBAs tooling that works exclusively in a language and environment that's familiar for them rather than splitting it across SQL and python/tcl/ruby/whatever. Not nearly every competent [relational] DBA is competent across multiple languages. Every competent [relational] DBA is competent in SQL.
Is it even possible to use set-based SQL to call EXEC SQL EXECUTE IMMEDIATE or sp_executesql on each statement in a set?
If you're matching a couple short strings, sure, don't bother overthinking the regex. If you're matching a lot of them, and/or they're long, then the extra time spent on making a single regex work will be worth it. The regex will work smarter than your hand-rolled code, and it also won't waste memory returning partial results.
Also: in my experience, almost every "disaster" regex comes from people not bothering to document and test what they write.
The lexer uses regexes but only for splitting the input stream of characters into tokens. Identifiers, integers, operators, strings, keywords, opening brackets and whatnot - each type of token is defined by a regex. This part is hopefully deterministic and simple, although the lexer matches regexes for all kinds of tokens at once, which is why lexer generators are often used to generate lexers.
The heavy lifting is done by the actual parser which tries to combine the tokens into something that makes sense from the point of the grammar.
So in this trick the sub-regexes between |'s define the tokens (the lexer part) while the group mechanism selects the single token that we want to keep (a very very simple parser).
I used to have a laminated sheet on my wall at an office because it was so terribly bad.
Named capture group: (?<name>foo)
Non-capturing group: (?:foo)
Lookahead: (?=foo)
For negative lookahead, change = to !: (?!foo)
For lookbehind, add <: (?<=foo)
For negative lookbehind, change = to !: (?<!foo)
(Not from memory, had to look everything up...)
right, great list but I'll forget it all by tomorrow.
You have to be careful about inputs like this though: “Inside a string”Tarzan”Again inside a string”
My attitude is generally that one should use regexes for matching regular languages and if one needs a stack or even Turing completeness then handle that in code around the regex.
What you want is something like all matches of regex("Tarzan") not contained in a match for regex("\"Tarzan\""), which is a bit trickier. That would require something like:
regex("Tarzan") - all-substrings(regex("\"Tarzan\""))
and I'm not sure regular languages are closed over the "all-substrings" operation. Actually I'm pretty sure they aren't.
I’ll take that as a complement.
Because the idea is to match Tarzan, but only if it is not preceded and followed by a quote.
Regex intersection and complement do not perform look-behind or trailing context.
Live demo:
This is the TXR Lisp interactive listener of TXR 265.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
TXR may be used in areas that are not necessarily well ventilated.
1> [#/Tarzan&~"Tarzan"/ "Jane shouted, \"Tarzan\""]
"Tarzan"
2> [#/Tarzan&~"Tarzan"/ "Jane shouted, \"Tarzan!\""]
"Tarzan"
The &~"Tarzan" makes absolutely no difference. The reason is that Tarzan matches exactly one string. The complement ~"Tarzan" matches a whole countable infinity of strings, and one of those is Tarzan. The intersection of that infinity and Tarzan is therefore Tarzan.Intersection with complement is useful like this:
Search for a three-character substring that is not cat:
3> [#/...&~cat/ "hat"]
"hat"
4> [#/...&~cat/ "dog"]
"dog"
5> [#/...&~cat/ "doggy"
"dog"
6> [#/...&~cat/ "cat"]
nil
7> [#/...&~cat/ "scatter"]
"sca"
8> [#/...&~cat/ "catalan"] ;; "cat" is skipped, then "ata" works.
"ata">we want to match Tarzan except when this exact word is in double-quotes
and so the reader might start thinking of ways to "match" this. The author then starts to mention ways to do this, but at the end, their trick is actually to not "match" it, but to remember it in a group. This will not match what the author says it will match, because if you do regex.test(string) it will return true when "Tarzan" appears, because it is in the or statement.
It appears the author is very good at storying-telling though.
$ perl -E'say q("Tarzan") =~ /"Tarzan"(*SKIP)(?!)|Tarzan/ '
$ perl -E'say q(Tarzan) =~ /"Tarzan"(*SKIP)(?!)|Tarzan/ '
1
$ printf '"Tarzan"' | pcre2grep '"Tarzan"(*SKIP)(?!)|Tarzan'
$ printf 'Tarzan' | pcre2grep '"Tarzan"(*SKIP)(?!)|Tarzan'
Tarzan> Before we proceed, I should point out some limitations of the technique:
The author clearly states you may have to add one or two extra line of code, in your case regex.test(string) may become `regex.match(string).group(1).length` > 0 or something along those lines.
The author explicitly states:
> so it will not work in a non-programming environment, such as a text editor's search-and-replace function or a grep command.
But in a programming environment, I will choose to have one more line of code over the extremely hard to read alternative regexes.
(...the trick is still cool, though; I can imagine other situations where it would be more useful. However it does seem like it potentially depends on the particular regex engine being used, in contrast to the author's claim about it being totally portable; yes, it'll compile on anything, but will it work?)
Lookbehind
Lookahead
Advanced handling of tags
Replace before matching
the best regex trick ever:
"Tarzan"|(Tarzan)
The whole site contains useful regex advice
Speaking of lengthy: this site breaks the iOS Safari scroll bar! It just disappears altogether (even when scrolling up or down to make it show, like you have to these days to please the UX designers in Palo Alto).
OK that's pretty clever (I certainly never thought of putting a capturing group inside only one side of an "or")...
...but it doesn't seem particularly useful? It probably won't work in most cases where this is just part of a larger expression. You're usually using capturing groups in a particular way for a good reason, and this would mess that up.
In contrast, the lookbehind+lookahead way is the "proper" and intuitive way to write it, and works as part of any larger expression.
So... +100 points for cleverness, but don't actually use this please. :)
I would say, the "proper" way is to have a separate line of code validating what's not there :)
(.)Tarzan(.)
Then in an additional line of code assert (Group 1 == Group 2) ≠ "
This shifts the logic out of regex and into the surrounding programming language context. That's arguably better, but the resulting regex is extremely dull and unclever.Maybe at this point you aren't using regex even. Nice, you solved two problems.
(I do appreciate regex and even use them a lot. But, I use them enough to avoid them as much as possible.)
(^|.)Tarzan(.|$)
Though I’m not 100% sure offhand what the result in the capturing groups would be.But generally, once you decide to use a regex in the first place, you might as well put as much regular everyday logic as you can in it. Otherwise you might as well look for "Tarzan" with a dumb string search.
Lookbehinds and lookaheads aren't rocket science. And you can always leave a comment about what they're doing if you're worried other team members won't grok the syntax.
Lookbehinds and lookaheads (especially negative lookbehinds) are rocket science.
What is "rocket science?" "Rocket science" is the feeling you get in math class where the instructor explains a proof to you in the clearest possible terms and you just don't get it. You have to listen to the explanation multiple times, preferably in a few different ways, and then you have to sleep on it, and then you get it, maybe.
But "rocket science" isn't just hard to understand. It's a hard problem where the consequences for failure are catastrophic. When you fail at rocket science, a multi-million dollar rocket explodes.
Anyone who's ever tried to teach lookbehinds to a newbie has seen it: you explain how lookbehinds work, and then ask the newbie to create a regex with negative lookbehind, to demonstrate mastery. I've done it a few times, and they never get it right, ever.
At best, they flub the syntax, but even once they get over that, they usually write the worst possible regex: a regex that works correctly on desired inputs but does the wrong thing on the input the regex is designed to reject.
This is a notorious problem with writing regexes, but it's way worse for negative lookbehind, because it's asserting that something isn't there, rather than querying for something that is there.
When I see a regex with negative lookbehind during code review, I ask for unit tests, not just comments. Reliably, regexes get even more complex when unit tests are added, because it's just so damn hard to write a correct regex with negative lookbeind.
I've never used the "trick" from TFA before, but it already sounds way easier to use than negative lookbehinds, and I'm curious to try it.
Things like greedy vs. non-greedy matching, matching newlines or not, handling Unicode correctly, inserting a capturing group when you actually needed a non-capturing group, making sure your regex works if it matches the start or end of a string, escaping characters -- those can be tricky.
On the other hand, lookaheads and lookbehinds are conceptually extremely straightforward, you just need a cheatsheet to remember the syntax is all.
Yes, this was sort of the idea as well (also see sibling response). I'd just as soon have 2 lines of code rather than a regex.
If anybody on your team doesn't understand regexes, you mean.
Save the clever stuff for where it's needed.
It's just not a particularly good "interface" for the task it is intended to achieve, a little more ability to be "verbose" at the possible price of succinctness I think would go a long way. I'm more-or-less waiting for the "blank" in: "blank" is to Python what Regex is to Perl.
"Find every 2nd instance of a dollar amount that is not encased in quotes" outputting <insert regex here> would be awesome
a. Code, once provided, can be broken down and understood at a far easier level than is required for composition;
b. Worst case, try several test cases to both increase comprehension and reduce the chance of 'gottcha's.
Shouldn't be too hard to stick with option 'a' as clear best practice, looking up any operators or syntax that aren't immediately obvious, the advantage being that the AI can use obscure tricks that you aren't initially aware of but you still have the opportunity to review and understand the regex, becoming better over time. It's theoretically auto-generated, but practically computer-assisted.
@P=split//,".URRUU\c8R";@d=split//,"\nrekcah xinU / lreP rehtona tsuJ";sub p{
@p{"r$p","u$p"}=(P,P);pipe"r$p","u$p";++$p;($q*=2)+=$f=!fork;map{$P=$P[$f^ord
($p{$_})&6];$p{$_}=/ ^$P/ix?$P:close$_}keys%p}p;p;p;p;p;map{$p{$_}=~/^[P.]/&&
close$_}%p;wait until$?;map{/^r/&&<$_>}%p;$_=$d[$q];sleep rand(2)if/\S/;printWhen I watched Idiocracy, a small optimist in me said "but surely the techies..." That optimist has died. We're fucked
This will sound like a forced joke but I genuinely didn't understand your phrase. I got stuck re-reading several times the "blank" in: "blank" part, but my mental language regex wasn't matching the expression.
I think the bug is caused by a bogus quote that causes a bad parameter expansion. My regex engine parses this better: the "blank" in: "blank is to Python what Regex is to Perl"
Off by one errors...
Parser combinators
- context sensitive matching
- matching with multi-char-exclusions
(regex is happy the most, when it's used to match "regular language" things)
So to search for words without any vowels just 'grep [^aeiou]'?
This style of writing is just obnoxious.
It's almost like nature, many simple rules coming together to make extremely clever and fairly complex ideas
not_this|(but_this)
... is interesting. But since it returns the match in a submatch I would say the \K approach is better: (?:not_this.*?)*\Kbut_this
Because usually when you try hard to accomplish something with a regex, you do not have the luxury to say "And then please disregard the match and look at the submatch instead".Now that I try it, it indeed does not work.
Maybe the author reads this and can look at it.
sadly, this trick still requires a code comment to explain. Python example:
# match tarzan but not "tarzan"
# see https://news.ycombinator.com/item?id=27774584
if "tarzan" == re.search(r'"tarzan"|(tarzan)', myvar)[1]:
...
which in practice means it probably deserves a function: if re_search_but_exclude(r'tarzan', myvar, '"tarzan"'):
...
I don't recommend monkeypatching re, i.e. re.search_but_exclude = ... if "tarzan" == re.search(r'"tarzan"|tarzan', myvar)[0]:
...
Am I missing something?foo = foo or [ 23, 42 ]
Or more generally:
foo = foo or ConstructSomeFoo()
If foo is None then it gets this default value or the newly constructed object, otherwise it's unchanged. Key here is that what's after "or" is not even evaluated if the first operand is already evaluating to True.
So, the left "Tarzan" eats up the matching substring that we do not want, while the right (Tarzan) matches what we do want, but only if the left one didn't already hit.
In abstract semantics of regular expressions, a|b and b|a are equivalent.
"Tarzan" will match one character earlier than Tarzan as the sting is scanned, so it would be discarded even if you flipped the order of the alternation.
This isn't true of examples where the good and bad marches can start at the same character.
Nice to see some examples of how ugly lookbehinds / lookaheads can be. And nice to have a new trick for avoiding them!
Although personally, I still think the most pragmatic solution in this case is usually to just filter out "Tarzan" values somewhere other than in regex.
I see non-look regexes as stepping over each of the characters where you can never go back in time - once a char is stepped over it is gone.
“Looks” allow you to step over characters to true|false match them, then step back in the string as if the look did not exist.
The Greatest Regex Trick Ever (2014) - https://news.ycombinator.com/item?id=10282121 - Sept 2015 (131 comments)
This one bugs me, because it's a cool enough trick, and I want to expand my thought process when it comes to regex.
The closest I can get visually in vim:
/"tarzan"\zs\|tarzan
(The \zs flag starts the cursor and highlighting at a given location inside of a larger regex match. I didn't use a capture group here because it didn't help.)Two problems:
1. This will still match the quoted word when pressing "n" but mostly unhighlights it. (see next point)
2. Whatever single character is after the unwanted match is highlighted, so this would only help for visually searching for reasonably long expressions.
-----
An alternative that I would use unless a special edge case was present (and this is basically the dumb version of the author's typical solutions):
/tarzan\ze[^"]
(\ze ends the match but continues to filter whatever follows)In a persistent edge case, I'd probably resort to macros or temporary replacement of the unwanted term. But that's not very satisfying, is it?
-----
More details:
Capture group references evidently work in vim's search mode. I hadn't tried until now. I only see utility in a few cases e.g. finding any duplicate word. The specific case given at the link does not work as-is. I'd need a way of evaluating the author's full expression and then only match the capture group. Is there a way to put the capture group \1 outside of the alternation?
There's possibly a way to use back-referencing or global search and execution or branches. The solution is also probably very clever and concise! I've tried a few permutations and am still stumped.
-----
Last best attempt:
/"\@<!tarzan"\@<!
(\@<! will match if the previous atom---in this case, double quotes---is not present.)An edge case where this falls apart? Single leading double quote e.g. "tarzan
Is there a better way?
Always over-document your regexes and assume people only have very basic regex skills.
This:
HEADER_PARAM =
/\s*[\w.]+=(?:[\w.]+|"(?:[^"\\]|\\.)*")?\s*/
is not as useful as this: /
\s* # Maybe whitespace at the beginning
[\w.]+ # Header key
= # Equals (yes:)
(?: # Header value from here
[\w.]+ # Just about anything
| # or
" # it could be wrapped in quotes
(?:
[^"\\] # Not quotes or backslashes
|
\\. # Sometimes escaped characters occur
)* # Could even be an empty string
"
)? # Maybe they didn't supply a value
\s*
/x
If you can use interpolation with your regexes you can extend this idea into (largely) self-documenting regexes.And this week I randomly thought about it.
I also thought I should have bookmarked it, because now I do not know where to find it again
/"Tarzan"|(Tarzan)/g
Probably one of my privacy plugins blocking something but I'm not going to debug someone else's page today.
I don't have a hard and fast rule of my own about regex complexity, but I do have a strong intuition over what's now ca. 25 years of working with regexes dating back to initial exposure in Perl 5 as a high schooler. That intuition boils down more or less to the idea that, when a regex grows too complex to comprehend at a glance, it's time to start thinking hard about replacing it with a proper parser, especially if it's operating over (as yet) imperfectly sanitized user input.
Sure, it's maybe a little more work up front, at least until you get good at writing small fast parsers - which doesn't take long, in my experience at least; formal training might make it easier still, but I've rarely felt the lack. In exchange for that small investment, you gain reliability and maintainability benefits throughout the lifetime of the code. Much of that comes from the simple source of no longer having to re-comprehend the hairball of punctuation that is any complex regex, before being able to modify it at all - something at which I was actually really good, as recently as a decade or so ago. The expertise has since expired through disuse, and that's given me no cause for regret; the thing about being a regex expert is that it's a really good skill for writing unreadable and subtly dangerous code, and not a skill good for much of anything else. Unreadable and subtly dangerous code was fine when I was a kid doing my own solo projects for fun, where the worst that'd happen is I might have to hit ^C. As an engineer on a team of engineers building software for production, it's not even something I would want to be good at doing.
You can get some surprisingly complex yet readable regexes in Perl by using qr//x[1] and decomposing the pieces into smaller qr//s that are then interpolated into the final pattern, along with proper inline comments in the regexes themselves.
Regexes are code.
Therefore, decomposition makes complex regexes both feel and actually be easier to reason about.
I do see a great opportunity to, by assuming interpolated qr// substrings have the locality the syntax falsely suggests, inadvertently create exactly that kind of mishap with it being minimally no easier, and potentially actually more difficult, to notice.
Write your code however you like, of course, including concatenating strings and passing the result to 'eval'. The last time I dealt with more Perl than a shell one-liner was around 2012, and that the language encourages this kind of thing is one of the reasons I'm glad of that.
And I use proper decomposition to keep it cognitively manageable. It's pretty clear that reasoning about composition is beyond you, but trust me that given two procedures that both do not have an undesirable property, one can rest assured that simple composition will not introduce that undesirable property.
The problem with RegEx is its "obscurity". However Maybe someone could write a nice testing tool that would throw millions of known exploits into each regex it finds in your code to see if it is vulnerable.
https://itnext.io/a-wild-way-to-check-if-a-number-is-prime-u...
The main trick for me was you first have to convert the number to unary, which was done outside of the regex.
was to convince programmers it didn't exist?
...because what could possibly go wrong? From the latest comment at the end of the page, the author would like you to know that the outcome is your problem, because you're using the wrong browser:
June 20, 2021 - 15:02
Subject: RE: Undoing whatever is hiding this page.
Hi Allen, try a different browser. There's no strange shading on the page, your browser is deciding to display it in a weird way. Regards, -Rex
Interestingly, the author also appears to control yu8.us
Breaking one's own content by https-ing one site but not another is a great example of why to not prop up a website's basic legibility on a third party dependency, even if it's one you own and control.
> Page copy protected against web site content infringement by Copyscape
(It's not that the annoyances aren't annoying, it's that they're so common that they lead to repetitive offtopicness that compounds into more boring threads.)
Good trick though.
Just make browsers for into “read only” mode where input cannot be accepted on non-secure pages. But don’t wall them out!
Parsing the tokenization result of a Dyck language still requires as context-free grammar.
It's not a badge of honor or a great trick to try that with regular expressions. It is using the wrong tool for the job.
Some of the candidates started by saying: "I know! I'll use a Regular Expression", to what I replied: "Great!, now you have TWO problems!"
^((1\d\d|2[0-4]\d|25[0-5]|[1-9]\d|[1-9])\.){3}(1\d\d|2[0-4]\d|25[0-5]|[1-9]\d|[1-9])$
Was the second problem "the interviewer"?No hire :-)
2134567890 is a valid IPv4 address
You can try and ping it
The only way in which
"Tarzan"|(Tarzan)
extracts a Tarzan that is not in quotes is when it is used for scanning the input for non-overlapping matches.(We know from lexical analysis with regexes, a form of non-overlapping extraction, that the "Tarzan" token is different from a Tarzan token. An identifier won't be recognized if it is in the middle of a string literal.)
It's not the regex itself, but a particular way of using it.
If the regex is used for finding all maximally long matching substrings, then it won't work. It will find "Tarzan" and it will find the Tarzan also within those quotes.
Notably, the regex will also fail if it is used to find a single match, like the leftmost. If the datum is a string like
"Tarzan", said Jane; Tarzan turned.
then the leftmost "Tarzan" will be found, and that's it. The regex will not find the leftmost Tarzan that is not wrapped in quotes.We cannot even use this to simply grep files for lines that have Tarzan that is not in quotes.
But I don't what environments return multiple matches from one evaluation.
"Tarzan", she said.
contains two matches for the regex "Tarzan"|Tarzan.
The first match is at character 0, for the "Tarzan" branch of the regex. The second match is at character 1 for the Tarzan branch of the regex.If matches can be overlapping, then the inner Tarzan is matched in spite of being surrounded in quotes, and the capture register is bound and all.
This works not as a property of the regex (what it matches), but the regex combined with a scanning algorithm that extracts non-overlapping matches.