Why Using .* in Regular Expressions Is Almost Never What You Actually Want
blog.mariusschulz.com
blog.mariusschulz.com
There is one legitimate use case of .* though: advancing to the last match of something. If you want to find the last digit in a string, /.*(\d)/ will readily find it for you.
http://www.charbase.com/2169-unicode-roman-numeral-ten
(even more in http://www.charbase.com/block/number-forms)
I'm not sure if Javascript matches these in its \d pattern, however, but I think that most regexp engines default to the ascii [0-9] unless you are using \p{Number}.
false
There's a modifier if you want to only match ASCII digits.
$ anchors to the end of the string, \D clears the non-digits from the end to allow \d to match the digit '2'.
In this case when finding the last match from the end, would the lazy quantifier reduce backtracking? e.g. /(\d)\D*?$/
On the two options generally:
/(\d)\D*$/
is problematic if you have a lot of digits, while /.*(\d)/
is problematic if you have a lot of text after the last digit. Both could potentially be optimized by the engine to run right-to-left (the former because it's anchored to the end and the latter because it greedily matches to the beginning), and then both would do well. I'm not sure if that happens in practice.Overall, I prefer the latter, both because I think it's clearer and because its perf characteristics hold up under a wider variety of inputs.
Edit: how do you make literal asterisks on HN without having a space after them?
/.*(\d)/
- Searches right-to-left, backtracking until it finds the match. /(\d)\D*$/
- Searches left-to-right, going forward a step, backtrack, forward, backtrack, until it finds a match.If you're looking for a match toward the end of a string, the .* version will be faster.
If you have a complicated regex $r, you can only negate it with (?:(?!$r).), and in that case, .$r is much easier to read :-)
I tend to avoid the non-greedy operator just because it often fails in terrible half-assed regex implementations (eg. visual studio 2010)
Arbitrary expressions can have arbitrary length, so excluding an expression simply will match it, fail the match, and backtrack to the next option.
(So there at least was an advantage)
"spotty" probably isn't the right word either, the change in 3.0 was to default to treating text as always being Unicode, the 'unicode' type in 2.x is reasonably complete (as these things go), just not the default treatment for text.
'acdha' (in the brother comment) wrote what I meant with "spotty" better and more pedagogical than I ever could. :-)
strasse = straße
or treating combining characters the same as their single character equivalents:
ñ = ñ
(That's LATIN SMALL LETTER N WITH TILDE and LATIN SMALL LETTER N followed by COMBINING TILDE)
A surprising number of languages (mostly everything but Perl) won't handle advanced uses like this.
The good news is that the next version of the stdlib regex module is being developed independently:
https://pypi.python.org/pypi/regex
Simply "pip install regex" and:
>>> regex.match(r"(?iV1)strasse", "stra\N{LATIN SMALL LETTER SHARP S}e").span()
(0, 6)
>>> regex.match(r"(?iV1)stra\N{LATIN SMALL LETTER SHARP S}e", "STRASSE").span()
(0, 7)I know this has been done forever, but usually only by extreme greybeards in Vi or Emacs world. The auto-highlighting now makes it possible for everyone to do it.
So that said, with the ability to restrict regexes to just a selection of text, it's more about regex golf--the fewest characters, the most productive--than it is about semantic correctness. If it works for my input, that's all that matters, because the regex is getting discarded thereafter.
Format a load of data using regexes first, then use it hard coded as string to do quick one off script to update the database. It beats trying to parse Excel directly, as you never know what data type a cell will return.
I like posts about details of software craftsmanship like this.
Edit: Cough, after rechecking... PCRE is not as universal as I thought. It seems I've been lucky. :-) http://en.wikipedia.org/wiki/Comparison_of_regular_expressio...
Edit 2: "Atomic groups" on that wikipedia link is when you can write a full grammar in a large regexp, right? Answer myself: No, it is the name for stopping backtracking. I've seen it as named "possessive" (perldoc perlre).
grep on the Mac used to use PCRE regexes if you used the -P option (`grep -P ....`), but beginning with OS X 10.8 the -P option was removed, so an important place that used to offer PCRE (default grep on a default Mac) actually removed support for it. They didn't replace it with something better; they just took it away. Maybe a Unicode issue?
Not only is PCRE not universal, overall support might even be waning.
Using an input string of abc123 he claims [a-z]+\d+ will match the entire string (which I agree with). He then says that [a-z]+?\d+? will only match abc1. Wouldn't it fail since the non-greedy match on [a-z] would just match 'a' causing the non-greedy match on \d to fail trying to match 'b'?
I used this tester posted elsewhere in the thread, it seems like since the lazy components expand "as needed" to achieve a match, it will succeed on "abc1".
EDIT: I wrapped it in a group for clarity.