Instead of
[ie]*-?[A-Z]+
it looks like they wanted [ie]?-?[A-Z]The harder ones I have dealt with are those looking for malformed syntax where the closing mark that might be missing could be several thousand characters after the opening mark, or the opening mark itself might be missing, across a data set that is several hundred million characters. So you need something very complex to find all the distinctive characteristics of the content that is supposed to be enclosed - while avoiding the many similar structures that give false positives. Sometimes the technically easier solution is too slow to run (look ahead and look behind, etc), so you need to pivot and use other regex features. It can take a day or two to get right.
https://stackoverflow.com/questions/1732348/regex-match-open...
Perfect
Regex to me, is pattern finding and abstraction taken to the extreme. I like the challenge
/
(?<=^|[ \t])
(?<currency_prefix_with_space>
(?<currency_prefix>
€|EUR|\$|USD
)
[ \t]?
)?
(?<number>
(?<integral>
-?
\d{1,3}
(?:[\.,]\d{3}|\d*)
)
[\.,]
(?<fraction>
\d{2,3}
)
)
(?<currency_postfix_with_space>
[ \t]?
(?<postfix_or_ending>
(?<currency_postfix>
€|EUR|\$|USD
)
) | (?<ending>
[ \t]|$|\n
)
)/xI would've probably approached it differently, trying to first get the 'inverted' match (i.e. ignore anything that isn't a currency-like pattern) and refine from there. A bit like this one I did a while back, to parse garbled strings that may occur after OCR [0]. I imagine the approach does not translate fully, because it's pattern extraction rather than validation.
The idea of an exclusionary approach sounds interesting as well. I'll have to think about that a bit.
Now I am enlightened.
We don’t need better language tools. Better parsers can, and already have, been implemented in libraries.
I’d have a look at parser combinators.
Is true though, or perhaps "state"? I know I had to come up with an algorithm because regexp alone couldn't do what I wanted (not even advanced features like lookahead, lookbehind, etc.)
Libraries like python's lrparsing [0] let you assign regex's (aka tokens) to variables, then build grammars by combining them using python expressions. For example:
while_statement = Keyword('while') + '(' + expression + ')' + block
statement = while_statement | ....
block = Token('{') + Repeat(statement + ';') + Token('}')
These grammars are more complex, with more rules you have to follow, but also more checking is done when they are compiled. They tend to mostly work once they do compile. So they are what you asked for, but there is no free lunch.On the down side, lrparsing is pure python so it's slower than python's inbuilt regex's.