Parsing JSON with a single regex
blogs.perl.org
blogs.perl.org
Perl regexes can be recursive: the operator $^R means "match here the whole regex itself", making the matching algorithm a full recursive descent parser. I suspect that parsing speed wouldn't be too bad because JSON is LL(1), so the parser shouldn't backtrack.
I don't think that the author is suggesting to use this in any useful code, he's just showing a cool hack. It kind of resonates well with the Perl aesthetics of short, "clever" code with as little as possible dependencies (and a fascination for regexes).
Would you like to call eval, recursion, macros, closures and many such things cheating?
Sorry but doing things without understanding or knowing the true consequences of what you are actually writing, always causes bugs. And that's true with any programming language.
Perl has many features which are designed to make the language powerful enough, to express solutions to very complex problems that other wise demands piles and piles of code in other languages. That's the whole point of the Perl language as a whole, you should take a really hard problem on a Friday night which you would find very difficult to code up in a weekend and then go to a finished script by Sunday evening.
Yes, in the context of "I did this with just a regexp".
> Sorry but doing things without understanding or knowing the true consequences of what you are actually writing, always causes bugs. And that's true with any programming language.
That's a cheap excuse for implementing this in a way that is just plain wrong, or too limited to be useful, when it could have been implemented correctly (at least if enough capable programmers still understood Perl's source code to work on it).
> Perl has many features which are designed to make the language powerful enough
I never disputed that. I used Perl almost exclusively in the past 14 years or so, so I know what I'm talking about. These broken language features are the biggest PITA about the language. They are useful for some very specific tasks, but useless for others where they could be used if they had been implemented correctly. Adding a feature with more gotchas than useful use cases is not the way to making a language more useful and enjoyale.
I agree with this in general, but not in this situation. Let's say you use a function you verify doesn't contain regex matching. Then someone else changes that function to contain regex - s/he shouldn't need to (and may not be able to) verify every single place where this function is used - introducing a hard to find bug. Who caused a bug here? In my opinion that's the regex engine/language creator who didn't decide to either handle this case correctly in all cases, or explicitly fail in this edge case.
(which is what I was expecting yeukhon's link to go to; it comes up in every discussion of abusing regular expressions to parse very non-regular languages, for good reason).
In practice, however, the regex engines of most programming languages have features that makes them capable of recognizing (a subset of?) type-2 grammars, hence the confusion.
[1] PHP manual for PCRE recursive patterns http://www.php.net/manual/en/regexp.reference.recursive.php
[2] Recursive subpatterns allow matching a language of the form (a * n)b(a * n) which is not regular.
Current day Perl regular expressions are far more powerful, and totally a very different beast.
They're more like general full-featured pattern matching engines for most forms of text.
Sadly, I couldn't find a way to access the nested matches, so it's quite pointless :)