There's never been an easier time to write your own language (2016)
joshsharp.com.au
joshsharp.com.au
And citing parsing isn't a great example. Parser generators have been around for ages. And they're usually not the hard part anyways. Defining a simple grammar and parsing it, even manually, isn't that terrible of a task. Getting decent error messages and figuring out recovery? That's trickier.
Code generation has certainly gotten easier. But you still need to go through the process of figuring out how to lower your abstractions. My language is still extremely basic but I've still had to map my high level types and control structures down to WebAssembly. LLVM won't do that for you.
There's also more that your average user expects if you want a language that people use. Decent tooling is important, so a language server and some syntax highlighting packages in different editors. Good error messages. Decent type inference. Most of these you can eschew in the first few iterations of your language but eventually you'll need them.
I feel bad criticizing this post because writing a language has been one of the most instructive experiences I've had. I've learned so much about code generation, typechecking, the WASM spec, etc. But it's still a lot of tough work to get to something people can use. I'm not sure parser generators and LLVM make it that much easier.
https://www.explainxkcd.com/wiki/index.php/2309:_X
(I don't try to troll the discussion)(it's just that XKCD is where many my mind runs to on some trigger words)
I also consider parsing and syntax to be the least interesting part of any language, as it is also a solved problem that requires little effort (modulo malformed code recovery) in both design and implementation when put in contrast with the rest of the compiler's functions and language design space.
1. Making "Algol 2020" is only interesting if you also throw person-decades of effort into making it competitive in production. I want to do other things with my time, therefore anything I attempt should not lead to that, which produces these other requirements to create a small yet useful language:
2. The implementation must leverage existing languages in a way that is uncomplicated to the user. Which means that it either compiles to some form of interpreter or to generated source code.
3. The language must focus on fully leveraging a specific data structure or family of data structures. What is and has long been fashionable in PL discussion is to elaborate upon symbolic expression. You do need some symbol definition to have a language, but the preferred orientation we use for many data structures is spatial: "top of stack", "bottom of tree", "traverse the graph", "loop over the array". Engaging with the vocabulary of the data in its context, and simply working to generalize upon that vocabulary and the bookkeeping it needs(iteration counters, selection markers, error cases etc.), rather than the generalities of algorithm definition, leads directly towards a tighter language. We can define many algorithms very well, and we're paying more attention to concurrency lately, but the software we're writing still mostly isn't about algorithms themselves. You don't end up writing one million lines of code because you have a mega-algorithm that's just really hard to express.
Consider regular expressions: a little string matching language, which can be usefully explained in a page or two. The idea of them has been around since the 50's, yet all the hip, popular languages today have implemented some syntax for them, making regex defacto one of the most common and long-lived programming languages in existence, outdoing "big" languages by many measures.
And ideally I'd like to engage in those terms: a language so small you don't really notice except to think "gee, that's handy."
Also, if this language isn't just going to be a toy, they need to write some sort of specification document about it. I'm tired of seeing new language websites announcing version 0.9 of the new language, accompanied by a statement that “we haven't yet written a language manual, but here are some example programs and some obsolete papers about prior versions”.
I used to work in a senior position at a multinational firm. At one point, I went to the site of a company that had been recently purchased by my employers. Their major product was a large online system that was substantially written in a custom language. At that point, since the principals in the original company had left, there was not one employee left who understood this language completely.
[0]: https://llvm.org/docs/tutorial/MyFirstLanguageFrontend/index...
But so does Diehl! It's no different!
He pulls together a 'disparate' parser library (actually he implements it himself, but it's a copy of an existing library) and he mentions using literally exactly the same backend (LLVM)!
What do you think the difference is that using a functional language makes?
It looks like exactly the same process to me.
Yes, if your language doesn’t fit C, your implementation could be more efficient (potentially a lot), but getting there isn’t easier.
I think we got a lot of complexity and options since then that distract from “writing your own language”.
I think a claim that we can do better than lex and yacc nowadays (e.g. using Haskell) would be a much stronger argument for “it’s much easier now”
Eg. Cowboy say "Hi there." Wait for 2 seconds. Cowboy run to Point P2 in 3 seconds.
It has been an eye opening and learning exercise. The advantage is even first time users can start using our language. Disadvantage natural language has such a wide variety of usage styles that it's a challenge.
And while you are at it, why not try to write a parser to reverse the process? :)
OTOH what makes a language usable is literally everything else - package manager, IDE support, compile/parsing.
Among other things, sure.
> OTOH what makes a language usable is literally everything else - package manager, IDE support, compile/parsing.
You have parsing/compiling in both categories. It's easier than ever to add IDE support for your homemade language, too, because of things like the LSP. Package management isn’t easy, but if you are building a specialized DSL and bundling the parts key to the domain, or if you are exposing a bridge to an existing ecosystem with strong library availability and decent package management, a language-specific package manager is probably not super important for usability.
But that could be because we haven't had any real innovation in languages for ages, so except for some superficialities and details, the languages are essentially the same.
Anybody that can do this things can probably take on any programming job, because they require a fundamental understanding of how computers and programs work.
Furthermore, It also help when suggesting realistic features to other language developers, or write proposals, or assess languages themselves.
Programming languages are not a solved problem (or parsing, or diagnostics, or...).
We are today, barely, getting practical solutions to how write safely, fast and ergonomic (for example, rust) and yet:
- Compiling performance is abysmal in most compiled langs (except pascal family)
- And errors...
- And debugging...
- And interactive compilers...
- And ... (thousands of other stuff)
Somebody must try to do this stuff, because if not, forever will be at mercy of C, C++, JS, the shell, unix, etc. Who wanna the cobol effect forever?
He also wrote “a plea for lean software”:
> “Abstract: Software's girth has surpassed its functionality, largely because hardware advances make this possible. The way to streamline software lies in disciplined methodologies and a return to the essentials. The paper discusses some causes of "fat software" and considers the Oberon system whose primary goal was to show that software can be developed with a fraction of the memory capacity and processor power usually required, without sacrificing flexibility, functionality, or user convenience.”
One recent article that touch it (for rust):
We’ve reached the ends of fast CPUs. To gain performance improvements, we need to use multiple cores, and process it in parallel.
I wonder if this is an opportunity for github or Microsoft to host a cloud compiler? Just check in your code to github, and the cloud compiler will compile it across thousands of servers, all running in parallel.
Cost: $100/month per engineer.
That could be a symptom but is not the root cause.Is just that the syntax/features choices (as choiced!) are hostile to fast performance.
Even going parallel, I bet a pascal compiler will beat it overly, because you can give rockets to turtles but are still turtles...
If anything, it's like writing a new novel. Why would anyone do that, when there are so many already written?
More crime novels alone than a person can read in a lifetime.
Must be approaching that for any given category or niche soon, if not already.
Are they named “novels” because they are novel? If so a novel experience is a distance from where you currently are, and can plausibly be had backwards in time by reading an old book as well as forwards in time reading a new book, or sideways reading a new genre.