Parsing: The Solved Problem That Isn't (2011)
tratt.net
tratt.net
They've pointed out that the difficulty of parsing, and in a sense our overconfidence that we can just code up parsers for random languages and input formats when we need them, is a pretty pervasive source of security bugs.
A lot of those bugs can occur when you have two different parsers that have a different notion of what language they're supposed to recognize, so it's possible to construct an input whose meaning the two parsers disagree on. That can have pretty serious ramifications if, for example, the first parser is deciding whether a requested action is authorized and the second parser is carrying out the action!
I'm kind of sad about this because I love whipping up regular expressions to extract data even from things that regular expressions technically can't handle correctly. But there's a good argument to be made that this habit is playing with fire much of the time, at least in systems that will end up handling untrusted input. And the Shellshock bug is a recent example of the way that your intuitions about whether your software will "handle untrusted input" in some use case can go out of date.
http://events.ccc.de/camp/2011/Fahrplan/events/4426.en.html
I went to that talk, and found it to be probably the most advanced and difficult math lecture I'd ever attended! (I was very grateful to see such mathematical sophistication brought to bear to protect PDF users, which is almost all of us.)
I learned a lot and PTAPG introduction specifically says that it avoids heavy theory. It has all algorithms required in very understandable presentation. It also says what this or that algorithm does to exploit that or this feature from grammar structure.
If you get through PL101 then picking up the stuff in the dragon book or any other book on parsing and compiling technology will be much easier.
Another resource I like is "Compiler Design: Virtual Machines" (http://smile.amazon.com/Compiler-Design-Machines-Reinhard-Wi...). Still going through that one but it is very readable and if you go through PL101 then you'll have all the tools to implement the virtual machines described in that book. It is much easier to write a compiler to target machine code or some other language like C when you've built a few targets yourself and understand the trade-offs involved.
There's also http://www.greatcodeclub.com/. I think one of the projects is a simple virtual machine and another one is a compiler. Well worth the admission price if you're a beginner and want some help getting started.
Hanselman's rule about finiteness of keystrokes applies and I recently wrote some notes about budding PL enthusiasts: http://www.scriptcrafty.com/tips-for-the-budding-pl-enthusia....
https://practicingruby.com/articles/parsing-json-the-hard-wa...
"Marpa is available as a open-source library. It is written in C, and the C library can be used directly or via a Perl interface."
http://jeffreykegler.github.io/Ocean-of-Awareness-blog/indiv...
I moved to PEG+Pratt exclusively and never needed anything beyond that, for even craziest grammars imaginable.
This is probably the best property of PEGs, they're extremely extensible and flexible, you can add features not possible in any other parsing technology, including high order parsers, dynamic extensibility, etc.
XL features 8 simple node types, 4 leafs (integer, real, text, name/symbol) and 4 inner nodes (infix, prefix, postfix and block). With that, you can use a rather standard looking syntax, yet have an inner parse tree structure that is practically as simple as Lisp. The scanner and parser together represent 1800 lines of C++ code with comments. In other words, with such a short parser, you have a language that reads like Python but has an abstract syntax tree that is barely more complicated than Lisp.
It's also the basis for the Tao 3D document description language used at Taodyne, so the approach has demonstrated its ability to work in an industrial project.
(prefix if 3)
instead of an error, since `if` is just a regular prefix operator. So, presumably there's some sort of well-formedness checking that goes on after parsing? Does anyone know more? The website says very little.Something like [A,B,C,D] writes as:
Block with [] containing Infix , containing Symbol A Infix , containing Symbol B Infix , containing Symbol C Symbol D
So this allows me to represent multifix along with their expected syntax. I could also represent for example:
[A, B; C, D]
if A then B else C unless D
for A in B..C loop D
The last one is actually one of the for loop forms in XL
The advantage of supporting multifix directly is that you get a more convenient and less confusing representation. For instance, given "if A then B else C" why should I expect (else (then (if A) B) C) and not (if (then A (else B C))? It's not obvious, and both are unwieldy compared to a ternary operator like (if/then/else A B C).
It's not simpler. XL started with a multifix representation, see http://mozart-dev.sourceforge.net/notes.html. Switching to the current representation was a MAJOR simplification.
The current representation captures the way humans parse the code. Infix captures "A+B" or "A and B". Prefix captures "+3" or "sin x". Postfix captures "3!" or "3km". Block captures "[A]", "(A)", "{A}" or indentation. Since humans perceive a difference, you need to record that difference somewhere. XL records that structure in the parse tree itself, not on side data structures such as grammar tables.
This approach also enables multiple tree shapes that overlap. Consider the following XL program (I replaced asterisks with slashes because the asterisk means "italics" for HN):
A/B+C -> multiply_and_add A, B, C
A+B/C -> multiply_and_add B, C, A
A+B -> add A, B
A*B -> mul A, B
In that case, I can match a multifix operator like multiply_and_add without needing a representation that would exclude matching A+B or A/B in other scenarios. This is especially important if you have type constraints on the arguments, e.g.:
A:matrix/B:matrix+C:matrix -> ...
A:matrix/B:real+C:real -> ...
Those would be checked against X/Y+Z in the code, but if X, Y and Z are real numbers, they would not match.
If you want to try it by yourself, you can download Tao from http://www.taodyne.com/shop/dev/en/content/10-compare-versio.... Tao uses XL as the basis of its dynamic document description. There's a tutorial here: http://www.taodyne.com/presentation/tutorial-2.0.html. Please note that the XL implementation used in Tao has several limitations, notably with local functions and closures.
Now, I'm not sure how you implemented multifix, but a good general implementation for it is not immediately obvious, so I suspect you simply used an algorithm that's more convoluted than necessary and made you overestimate the scheme's actual complexity. Implemented properly, though, it's ridiculously simple. To demonstrate, I have hacked together a JavaScript version here (sorry, I didn't get around to commenting the code yet):
http://jsfiddle.net/ed37wy5k/2/
The core of the parser is the oparse function, which clocks in at 32 lines of code. Along with a basic tokenizer, order function, evaluator and some fancy display code, the whole thing is barely over 200 LOC. I dug around for XL's source and from what I can tell, what you have is not simpler than this.
I also never suggested making multifix operators like multiply_and_add. Multifix is for operators which are inherently ternary, quaternary and so forth. For instance, when I see "a ? b : c" I don't parse it as prefix, postfix or binary infix, I parse it as ternary.
Underneath, there is a rewrite operator, ->, that means to transform the shape on the left into the shape on the right.
The precise semantics are a tiny bit more complicated than that, but it iterates until there is no more transformation that applies. Some specific rewrites will be transformed into "opcodes" of the underlying machine model. So A+B with integers on each side will execute an addition (the inner representation of the transform is something like A+B -> opcode Add).
For more details, you can read http://xlr.sourceforge.net/sites/default/files/XLRef.pdf. I badly need to update it, it's out of date, I have a more recent version, but I need to learn the new way to connect to SourceForge.
I did simplify the AST further than XL and added a few invariants. I have 3 leafs (literal, symbol, void) and 3 inner (send, multi, data). The inners roughly work like this:
f x (send f x)
[x] x
[x, y, ...] (multi x y ...)
{x, y, ...} (data x y ...)
a + b (send + (data a b))
+ a (send + (data (void) a))
a + (send + (data a (void)))
[] serve as grouping brackets and are therefore purposefully erased from the parse tree if they only contain one expression. Prefix/postfix operators are interpreted as infix operators that have a blank operand on either side. Infix operators are themselves reduced to send/data, so "a + b" is equivalent to "[+]{a, b}". I find that a bit more robust and easier to manipulate generically.When you look at the simplicity of OMeta[1] you feel the difference. But I am not aware of any (production ready) OMeta implementation for C++.
It's all about descent parsing and memoisation (thus the name).
And actually, my experience with packrat parsers (mostly in ruby) in other languages has been that they actually slow things down on moderately or more complex grammars by massively exploding memory use and thus allocation pressures. Turning it off can make it faster, especially on complex grammars. It's a pretty good case study in how optimizing an O(n^2) worst case to O(n) does not always improve things.
That said, I'm not against the principle, but the shotgun approach to it can be brutally bad. You really only want to memoize the paths that are actually likely to backtrack. Or simple grammars where O(n^2) memory use is not going to balloon your memory use too much it's a clear win.
Of course not, see the link I posted. PEGs are TDPL and recursive descent (not the other way around), and an alternative to CFGs. PEGs were coined by Bryan Ford, who then coined packrat parsing based on them.
You want to do smarter things than just memoising everything, yeah. Especially for complex grammars.
See http://ialab.cs.tsukuba.ac.jp/~mizusima/publications/paste51...
I noticed with the languages I was working on, the problems could be resolved by being smarter with the lookahead: this parser allows for context-free lookahead matching to resolve (or detect and defer) ambiguities.
That makes it possible to do neat things like parse C snippets without full type information or deal with keywords that aren't always keywords (eg, await in C#).
If you want to study grammars in an abstract sense, then think of them this way, and that's fine. If you want to build a parser for a programming language, don't use any of this stuff. Just write code to parse the language in a straightforward way. You'll get a lot more done and the resulting system will be much nicer for you and your users.
Your parser will probably not end up being able to handle the problems outlined in the article unless you take that theory into account before starting to program.
This approach is why many consider parsing to be a solved problem, so it's certainly a valid approach. However, it's not the only valid approach.
For example, "straightforward" parsers often give terrible error messages: when the intended branch (eg. if/then/else) fails, the parser will backtrack and try a more general alternative (eg. a function call). Not only does this give an incorrect error (eg. "no such function 'esle'"), but it might actually succeed! In which case, the parser will be in the wrong state to parse the following text, and gives a non-sensical message (eg. "unexpected '('" several lines later).
This is an important problem, since these messages can only be decyphered by those who know enough about the syntax to avoid hitting them very often! Inexperienced users will see error messages over and over again, and have no idea that they're being asked to fix non-existent errors in incorrect positions.
ADDENDUM: In particular, the article mentions PEGs. Not in an unambiguously positive light, but nonetheless: do you count PEGs as being a useless overly academic tool?
The LANGSEC project has done (in my opinion) a pretty good job arguing that this behavior causes many security issues in software. Perhaps only to some extent for programming languages in general, but in particular for data format languages and network protocols. The issue is slightly mitigated for programming languages because they typically don't have many wildly varying parsers; but (programming) languages without clearly specified grammars certainly have had problems with this in the past.
Are you familiar with their work? Do you just disagree with their conclusions?
[LANGSEC]: http://langsec.org