Regexes: The Bad, the Better, and the Best
loggly.com
loggly.com
I'd rather word this as "more specific is better". Like say for a U.S. phone number (minus area code for simplicity),
\d{3}-?\d{4}
is better than .*-?.*
because it's more specific."Longer is better" is only useful for helping identify which regex is better, not for helping write better regexes.
"In most flavors that support Unicode, \d includes all digits from all scripts." [1]
EDIT: But I guess for the majority of use cases it doesn't matter [2] since PCRE is the norm almost everywhere.
[1] http://www.regular-expressions.info/shorthand.html
[2] "Notable exceptions are Java, JavaScript, and PCRE. These Unicode flavors match only ASCII digits with \d"
[0-9]{3}-?[0-9]{4}
is even more specific (and faster) because \d would match other digit characters outside of [0-9].[0-9][0-9][0-9]-?[0-9][0-9][0-9][0-9]
In most places I used regexes, efficiency is the least of my concerns, though for a company whose product is focused around parsing massive amounts of logs, focusing on performance does make sense.
/.*? (.*?)\[(.*?)\]:.*/If you are searching a very large file for a very few occurrences of the expected match then this optimization is not so bad.
If you are running line-by-line through a very large log file to extract just those two pieces of information per line, then throw away the first N characters in each line (where N is the hopefully-constant length of your timestamps plus that space char) and start the regex engine at the beginning of the expected match. Then it doesn't have to waste any time passing over those chars.
Even if the exact details above aren't quite right the principal is (and is well-known): Avoid premature optimization! (And the corollary: Measure it. Profile your code, don't guess, you're probably wrong.)
AFAIK the only feature of regexes that require backtracking are back-references, as long as your regex doesn't use it why doesn't PCRE switch to the more efficient algorithm, and use the backtracking algorithm only if you actually need the feature that requires backtracking?
┌┐
↓│
┌────┴┐
┌─→│ │←──┐
│ └─────┘ │
╔════╧═╗ ┌──┴──┐
├─→║ ╟───A──→│ │
╚══════╝ └─┬───┘
↑ │
└─────A,B────┘
Three states if you want to process the whole input. If your model is "reject when you fail to find an appropriate transition" rather than "reject if, after processing the string, you're in a reject state", then you don't need the failure trap and you can do it in two states.Backtracking is definitely not required, nor helpful.
See https://swtch.com/~rsc/regexp/ for information on finite-state-machine implementations of regexp matching.
What? Why is that? If the NFA has n states, then the DFA in principle might need one state for every possible set of states the NFA might be in, of which there are 2^n. Where does n^2 come from?
https://metacpan.org/pod/Regexp::Debugger
After installing the module it can be started simply by calling `rxrx` on the command line.
As it stands, the 'good' regex will look through the whole line for any occurrence of 1 or 2 in the line, then start applying the rest of the regex from there. Putting an anchor in means it only looks at the first character for that 1 or 2.
Does anyone know a good browser-based game that trains you to do regexes well?