HNHacker News
TopNewBestAskShowJobs

renjipanicker

15 karma · joined July 24, 2011

submissionscomments
renjipanicker··on Show HN: Yantra – an LALR(1) parser generator for C++
Not currently on a concrete roadmap, but the architecture was deliberately left open for this. The C++ code generator is literally named cpp_generator.cpp, not just generator.cpp, specifically leaving room for other backends later.

One clarification on the "low-level speed from Python" part though: I think the more useful path there isn't a separate Python code-generator (which would mean Python itself doing the AST walking, probably giving up most of that speed), but Python bindings over the generated C++ classes, something like pybind11 or nanobind wrapping the AST and walker types. That gets you both: parse and walk at C++ speed, call it from Python. A true Python backend would be a different, slower thing aimed at a different use case (prototyping grammars without a C++ toolchain, say).

No promises on timeline, but if there's real interest I'd probably go the bindings route first.

renjipanicker··on Show HN: Yantra – an LALR(1) parser generator for C++
Mostly still true, yeah. A few tools have made real progress on error tolerance, continuing past a mistake rather than just stopping, tree-sitter is probably the best example, it's explicitly built to produce a best-effort tree from broken input, which is why it's good for editors. ANTLR also has configurable recovery strategies, single-token insertion/deletion heuristics and the like.

But "tolerant" isn't the same as "as good as hand-written." A hand-rolled recursive descent parser can say something like "missing semicolon after return statement" because the code knows exactly what construct it's in.

A generated parser's error is usually derived mechanically from the state machine, "expected one of: X, Y, Z, got W", which is correct but generic. This is what Yantra does at the moment. Closing that specific gap would mostly require hand-authored, context-specific messages layered on top. But its a good problem to solve.

For yantra specifically, it doesn't have error recovery at all yet. A syntax or lexer error just stops parsing at that point, no resynchronization, no continuing to find more errors in one pass. It's a known, documented gap, not something I'd claim is solved. For the kind of smaller or evolving DSLs this is aimed at, that's probably an acceptable tradeoff, but it does exist as a limitation.

renjipanicker··on Show HN: Yantra – an LALR(1) parser generator for C++
Thanks, I really appreciate that. The README's Quick Start should get you to a working parser in a couple minutes, and I'm around if you need any assistance.
renjipanicker··on Show HN: Yantra – an LALR(1) parser generator for C++
Fair and true. Most production compilers (Clang, rustc, Go) have moved to hand-written recursive descent, largely for error messages and debuggability: a hand-rolled parser can say exactly what went wrong and try to recover, a generated one is working from a state table.

One thing worth separating out though: recursive descent is already top-down, so it gets "build the tree, then decide what to do with it" for free, the same way ANTLR's LL(*) does. The intresting part of what I built is that it's getting that capability while keeping LALR's bottom-up table-driven parsing.

Where a grammar-first tool like this is more useful is smaller or evolving DSLs, where you want the grammar as a readable, declarative spec with automatic conflict detection instead of hand-tuned lookahead logic, and cases like generating multiple outputs (e.g. C++ and Java) from one grammar, which is awkward to bolt onto a hand-rolled parser after the fact.

renjipanicker··on Show HN: Yantra – an LALR(1) parser generator for C++
Right, mid-rule actions genuinely let you shuttle state into a child before it's parsed, so it's not nothing.

But as you said, it's still one linear left-to-right pass. Appreciate you drawing the line precisely.

renjipanicker··on Is C/C++ worth it?
While most of these discussions tend to focus on performance benchmarks as the primary criteria, there is another factor to consider.

To a large extent, it depends on the target platform. If you developing an application to work across the widest possible range of desktop and mobile platforms, C++ is your best bet.

Using C++, you can have a single code base that can be compiled for Windows, OSX, Linux, iOS, Android and WinMo. There is no other language that gives you that.

The UI code is not portable in any case, but that would be a common problem regardless of language.