HNHacker News
TopNewBestAskShowJobs

overthenexttwod

2 karma · joined August 3, 2026

submissionscomments
overthenexttwod··on Parsing Expression Grammar vs. Regexes: Building Org Parser in Lisp, Export HTML
Over the years, I have implemented parsers numerous times, both at work and for side projects, so writing recursive-descent parsers from scratch has become second nature. Once you understand the mechanics, writing a parser by hand is straightforward and offers distinct advantages, particularly much greater flexibility with error handling and reporting. Because of that, I had always viewed PEGs and parser generators as tools primarily for people who couldn’t hand-roll their own because of their circumstances or skill level. (If you look at projects that are neither understaffed nor underskilled you'll find that hand-written recursive-descent parsers are very common: Clang, Go, Rust, TypeScript, Swift and Lua all have hand-written parsers.)

LLMs have changed the equation. It used to take me 2-3 hours to write a parser for a moderately complex grammar. Now, if I hand an LLM a loosely written, BNF-ish grammar and ask for a recursive-descent parser, it finishes the job in five minutes. At this point, writing them by hand is hard to justify. The model does it substantially faster, and lately, often better than I do.

Which makes me wonder: what’s the appeal of PEGs or parser generators now? They used to make sense when hand-writing wasn't practical, but what compelling reasons are left to use them today?

overthenexttwod··on Prevent cognitive debt by manually retyping LLM-generated code
I share the author's sentiment. I also think it's important to fully understand a codebase I own. So much is naturally lost when you let LLMs generate code for you, and the mental model of what the added code does is one of the biggest losses. When you write code by hand, you build that model as you go, and it's enormously helpful later - when adding a feature, or when debugging behavior you didn't expect.

I've been trying to address this by telling LLMs primarily how the code should be structured, rather than only what it should do. Still, any design I hand to an LLM will be underspecified in one way or another (if it were fully specified, it would just be code), and the LLM fills those gaps somehow - which adds to the cognitive debt, slowly but surely.

Retyping LLM-generated code is an interesting solution. You'd certainly end up understanding the generated code better than if you merely reviewed it, but I doubt it produces a mental model as reliable as the one you'd build writing the code yourself. The longer you think, the better your mental model gets - and outsourcing the thinking to an LLM means thinking less.

That said, I've started to wonder whether I'm solving the wrong problem. Should I really insist on an accurate mental model of the code I own? We'd find it strange for a non-engineering manager to try to fully understand every piece of code their reports produce. If that's the right analogy, then as LLMs' agentic capabilities improve, maybe we should stop treating LLMs as tools that boost our own productivity as a software engineer and start treating them as independent agents we manage and steer.