Alternatives to Regular Expressions
c2.com
c2.com
Somebody please tell me this is sarcasm...
The best summary of the problems I’ve seen is from Larry Wall, highly recommended if you haven’t read it before.
First half of this page: http://perl6.org/archive/doc/design/apo/A05.html
Unfortunately, I think at this point regular expressions are too firmly entrenched in too many places to properly replace.
> In fact, regular expression culture is a mess, and I share some of the blame for making it that way. Since my mother always told me to clean up my own messes, I suppose I'll have to do just that.
So Larry aimed to do better in Perl 6. Do you think he is as good at solving problems as he is at identifying them -- or more specifically, what do you think of the "Rules" system[1] he came up with for Perl 6?
It's a mini language, and like any language it can be clear or opaque in its meaning depending on the skill and intention of the creator.
I.e. it's very easy to write something like /https?:\/\/www.foo.com/ (e.g. you want to whitelist your own domain for some action) and forget that this also matches wwwxfoo.com which could be owned by anyone (it's also a common mistake to forget the anchors, in this case evil.com/?https://www.foo.com might also match).
It is pretty easy to come up with a better syntax, but it's going to be hard to convince everyone to use it. As someone else commented on this thread, the current way of doing things is pretty entrenched.
http://www.rebol.net/wiki/Common_Parse_Patterns
https://en.wikibooks.org/wiki/REBOL_Programming/Language_Fea...
http://blog.hostilefork.com/why-rebol-red-parse-cool/
http://rebol-land.blogspot.co.uk/2013/03/rebols-answer-to-re...
1. Regular expressions are often used instead of existing parsers. XML, CSV, file paths, URIs etc. all already have fast, well tested and correct parsers.
2. Often the thing worked on should not be a string in the first place. For example comma separated string is used instead of a list and regular expression emulate list operations.
3. Some things could be processed much easier with different tool - for example recursive descent parser. Yet the developer still tries to parse arithmetic expressions with a hammer.
4. They are often developed via trial and error. They are either first thing which worked for a simple case or are 10 line monsters riddled with exceptions from exceptions.
5. They are often part of hacks and workarounds. For example User-Agent is matched to work around bugs in browsers.
3. They give very limited feedback to user. There is either a match or no match. There is no way to tell what and where is broken.
There are valid use cases for regular expressions. It's even possible to write correct code with them. It's just a rare sight.
As DOM query:
document.getElementsByClassName("title")[0].parentElement.getElementsByTagName("a")[1].href
This will break:* When title element no longer has "title" class.
* When title is no longer a sibling of link.
* When link is no longer 2nd link of its parent.
As regular expression:
document.documentElement.innerHTML.match('td class="title">.*a href="([^"]*)"')[1]
This will break:* On any white space change.
* On any new attributes on td or a.
* When ' is used instead of "
* When href includes escaped "
* In most cases when DOM query will break.
Many of those can happen without any server-side changes. It will sometimes works sometimes won't - making it hard to test.
There are cases when regular expression will break less often than DOM but DOM is easier to reason about, more predictable and has less corner cases.
1. Those hosts are still broken. The next person will have to jump the same hoops to support them.
2. You parser is very permissive. It will encourage people to create even more broken implementations.
3. The specification of this protocol is now worthless. There is no way to safely add new functionality. Any new element or attribute can break those regexes. Everyone has to take every implementation into account.
4. You are probably missing some corner cases like CDATA elements or quoted characters.
From your point of view, it probably makes sense to support even broken sites. But, you are helping to create next HTML - where every implementation works differently and you have to test everything on every browser.
(also, too many kids on my lawn etc)
People don't want to learn because it's hard?
"Most users of regexps or any other device do not need to read a thick theory book on it."
Which is why they should at the very least read Friedl.
"Cobbling together expressions from online tutorials is a perfectly good way to learn effectively."
No it's not, otherwise we wouldn't have so many bad regexes and people asking silly questions about them, would we?
"Admittedly, it is not the best way to write production code, but there are often economic considerations that override the expert's desire for perfect design and implementation."
Nobody's talking about 'perfect design and implementation', just 'not be a moron' level. Because car analogies are everbody's favorite, let me use one here: no-one is saying that one should have a Formula 1 licence before driving (using regular expressions); just to have more than 20% vision in both eyes and not to drive after drinking 5 beers while texting your wife that you're on your way. Which is the car-driving equivalent of the majority of regex uses in the wild.
There is truth in that, but I meant to address your comment about other people's lack of education. Education (certificate granted by teaching authority), how good soever it may be, is not necessary to have knowledge or skill. Similarly, reading Friedl how good soever his work is, is not necessary to learn or use regular expressions.
http://doc.perl6.org/language/regexes http://doc.perl6.org/language/grammars
No. If you're even thinking about defining a language in XML you've probably already screwed up some place.
http://lua-users.org/wiki/PatternsTutorial
At first I found it a bit confusion but they're actually pretty great. For me at least 90% of the tasks I want to do with regex can be done with scanf, and Lua's little extension covers the remaining 10% quite well.
I know lots of people often say that it would be nicer to use functional composition for regex instead of strings because strings are too confusing, but I disagree with this. The confusion of regex to me is not from the string representation, it is that some characters are "special" while others are "normal" (including whitespace). At first it appears that most characters are "normal" and so you can start from some example and generalize the string until it matches all the things you want - but once you start putting parenthesis and such in you start to realize that most of the string wont be matched "normally" and it is better to start thinking like a grammar and write it from scratch. This double thinking is pretty annoying.
For this reason, to me the Lua patterns are really the only alternative I've come across to regex that I've liked. They've got nice compact expressive syntax, can really easily do most of the matching tasks I need due to the scanf base, (almost) all the "special" characters begin with %, and the complex cases can still be matched.
There is still some backtracking behaviour that means you can create expressions that are slow, but for most use cases it is not an issue.
_That_ is a bug breeding ground.
Now, if what you are after are regular expressions with a different syntax.... well, maybe you're onto something here. But I would still call it a regex, personally.
That's an alternative.
Its not the only alternative. Monadic Parser Combinators are a thing, after all.
Idea is that it can build multiple layers of matched items. In this example, layer 0 is "pattern lines" then on top of this, gets mapped tokens representation at "parttern tokens".
and lastly, over lines and tokens goes AST objects, which are represented by language code.
Hierarchy and context-based patterns referencing and construction can be easily achieved.
Ideally, I would like to build this with realtime update feature, when used in source code editor. In IDE, when I write some characters in code, these changes propagate into parsed layers objects and updates what is needed.
StringReplace["Mad Hatter", "M"|"H"~~a_~~b_..:>"B"<>a<>"g"]
returns "Bag Bager"
The nice thing is that a string expression can use a RegularExpression as part of the pattern, but it's often not neccesary.
Nothing in this post is a nice thing :)
> Alphanumeric characters and the underscore _ are literal matches. All other characters must either be escaped with a backslash (for example \: to match a colon), or included in quotes.
k/q also has pattern matching.
Personally, I get more mileage out of BRE than anything else. Simple and effective.