Is it a must for every programmer to learn regular expressions?
programmers.stackexchange.com
programmers.stackexchange.com
But it's also a must that you realize that despite the fact that they're exceptionally useful and widely supported, regexes are a disgusting abomination, one that we should be absolutely mortified to be associated with. It's one of the worst syntaxes to ever be invented, and every one of us should feel the cold stink of UX failure wash over us every time we write a regex. If we ever catch ourselves writing a DSL that in any way, shape, or form resembles regular expressions, we should stop immediately and ask what the fuck is wrong with us and why we're being so opaque and random. Regular expressions are quite literally one of the worst syntaxes to ever be introduced in our field.
I worry a lot when someone doesn't know regular expressions at all. But I worry far more when someone thinks they're beautiful. That person has far too high a tolerance for unintuitive syntax and code, and will cause vastly more damage to my codebase than even the rank amateur that still uses "goto" on a regular basis.
Which is not to say that we shouldn't lean on regexes heavily anyways when they're appropriate - as programmers our primary job description is that we're paid good money to work with shitty interfaces in order to express simple ideas and algorithms.
Okay, so they're ugly and hard to read at first glance. Considering the purpose of a regex, I can't think of another way to implement them that doesn't involve typing more characters needlessly (therefore even making it more hard to comprehend).
\b[A-Z0-9._%-]+@[A-Z0-9.-]+\.[A-Z]{2,4}\b
Which is also what happens when a cat walks across the keyboard. if "@" in email and "." in email.split("@")[1]:
send_verification(email)
...but you should probably also check for common misspellings like "gmial.com" etc.[1] http://en.wikipedia.org/wiki/Email_address#Syntax [2] http://en.wikipedia.org/wiki/List_of_Internet_top-level_doma...
(?x: #standard token is an uppercase letter, digit, dot, or hyphen
\b
[A-Z0-9._%-]+ #1 or more of standard token, underscore, or percent sign
@ #at-sign
[A-Z0-9.-]+ #1 or more standard tokens
\. #dot
[A-Z]{2,4} #2 to 4 uppercase letters
\b
)
We can easily optimize for readability with regex syntax.[It's also perfectly obvious that if this is an attempt to match email addresses that it's not a very good one - but I don't know the context where it's supposed to be used, it might be good enough for whatever the author intended].
The greatest syntactical atrocity in regexes is that they don't have the `x` modifier (in Perl parlance) on by default. This means that you can't use whitespace to chunk code into meaningful bits, nor can you comment it to easily document what does what or explain a particularly hairy section to handle some weird edge case. This means that regexes degenerate a lot faster than ordinary code in terms of readability.
Edit: Misplaced close paren.
I always wonder what regexes would look like if they were derived from Python instead.
It is true that Perl reformed the syntax in important ways (to the better, if you ask me), and later on extended it a lot, but it's certainly not a Perl invention.
Anything that comes out of the whole chomskian hierarchy stuff isn't going to look intuitive. But the point is that it is a particular, very rigorously defined system of representation. And various systems of representation are always more or less intuitively accessible - and come with a whole set of trade offs around what they can represent vs their ease of use etc...
These things just are - written into the laws of the world. They are discovered - not invented. We pick them up and use them as we would rocks left lying around. We find better ones when we can and fashion better ones when we can too...
Why not? Type-0 (recursively enumerable) languages are equivalent to Turing machines, and we've managed to invent some pretty good syntaxes for that. The main problem with regexes really is the syntax. Regular languages are a lot easier to understand (IMO) if you look at the left/right-linear grammars that define the same language as the regex.
Regex syntax as a representation is very close to the FSA used for matching, and that's not necessarily the representation best suited for human consumption.
This applies only to the theoretical computer science regexes. The practical regexes are very different in this aspect.
Practical regexes are neither discovered nor invented - they are constructed. The theoretical regexes are just the basis but on top of that there's a lot of features added. Some of them are just syntactic sugar for theoretical regexes but others actually make the language non-regular.
Groups that match what previous named groups have matched are definitely in this category.
I think the real problem is that we lack (or don't learn) good tools to bridge the gap between regular expressions and 'custom parser'. We're reluctant to refactor from '1 line of just-starting-to-be-horrible regex' to tens or hundreds of lines (depending on language and libraries) to do it 'properly', and so we end up stretching regular expressions beyond the point where they make life easier.
Perl has Parse::RecDescent (and probably several others), which is pretty close to the right thing, and clearly it's very doable in a lot of languages - anyone got any suggestions in other languages?
Of course, you would also have to be able to nest these for more advanced matching...
Is there something like this already in existence? :)
> Regular expressions are a tool. It happens to be a
> very useful tool, so many people choose to learn how
> to use it. However, there's no "requirement" for
> you to learn how to use this particular tool, any
> more than there is a "requirement" for you to learn
> anything else.
Nails it. I do think most programmers will eventually run across a problem to which the solution
is 'Use Regex', but it's not an absolute "must" like boolean logic.[1] http://programmers.stackexchange.com/questions/133968/is-it-...
"Regular expressions are hard to write, hard to write well, and can be expensive relative to other technologies... Standard lexing and parsing techniques are so easy to write, so general, and so adaptable there's no reason to use regular expressions.
"Another way to look at it is that lexers and parsing are matching statically-defined patterns, but regular expressions' strength is that they provide a way to express patterns dynamically. They're great in text editors and search tools, but when you know at compile time whatall the things are you're looking for, regular expressions bring far more generality and flexibility than you need.
"Encouraging regular expressions as a panacea for all text processing problems is not only lazy and poor engineering, it also reinforces their use by people who shouldn't be using them at all."
http://commandcenter.blogspot.com/2011/08/regular-expression...
http://news.ycombinator.com/item?id=2915137
Personally, I think they're pretty darn useful and too powerful to not learn, but Pike's comment makes me think that maybe they're also a crutch that I've relied on too much rather than learning enough about lexing/parsing.
Also if you are familiar with this quote:
"Some people, when confronted with a problem, think
'I know, I'll use regular expressions.' Now they have two problems."
Read Jeffrey Friedl's research into the quote - http://regex.info/blog/2006-09-15/247I would say it's a must to take your coding skills to the next level.
If you ask "do you personally really need regexps" I'll tell you: don't learn them. As you're asking that question at all, I understand that you're not interested to learn and that you are looking for an excuse not to learn, so do something that interests you.
Surprisingly, here: http://www.regexbuddy.com/
The help file alone will make you want to buy it.
However, I try to avoid using them in my code unless they improve readability. Using re.VERBOSE can help in Python.
If you find a regex online, you should definitely reference it in your code, to help provide background understanding, such as validating a UK Post Code.
I reach for refex frequently but almost never as something I add to the code I'm writing. They're just amazing for filtering and bulk editing text.
[1] Thinking of embedded systems developers creating code for, say, automotive entertainment systems, or control code for scientific or medical hardware.
Don't they have log files to read?
But the point is I can't imagine someone getting to that point without previously having learned it.
It helps to know the basic syntax by memory, but you can just as easily look it up if you understand how they work.
[i]
Many don't know the difference nor implications of .* ? vs .*And on and on. It amazes me that these people actually code for a living.
But you will be limited. You will run into a lot of situations where your ignorance will create friction and limitations. Something as simple as editing an nginx configuration file, for example.
You should learn regexes, you should be able to use a regex as necessary and be able to understand regexes you come across, you should understand when they are most useful, and when they are not. You should learn their limits and your limits in using them as well. When used appropriately they can be a potent addition to your technical knowledge. When they should be used but are avoided the result is typically a massive inflation of the effort to solve a simple problem. And when used incorrectly they can result in an inflation of intractable complexity (much like any technology).