So you want to design a programming language (2017)
cs.lmu.edu
cs.lmu.edu
References 1. https://docs.racket-lang.org/guide/languages.html 2. https://docs.racket-lang.org/turnstile/
Examples 1. https://github.com/soegaard/urlang 2. https://github.com/rjnw/sham 3. https://github.com/ShawSumma/lure 4. https://github.com/racket/rhombus-prototype 5. https://github.com/lexi-lambda/hackett
I think the article is more aimed at the latter, whereas if you just want to get something working and have fun, your approach makes more sense.
The list of prerequisites in the article is extensive, overwhelming, and largely unnecessary if you're just getting started. Most of the concepts that it lists as "prerequisites" are ones that I finally understood while implementing a toy language, not stuff that I studied ahead of time.
I don't think I would ever want to do anything beyond a DSL or bytecode for tiny 128 byte ish embedded config or the like, and even then only if nothing else worked.
The odds of usefulness are low even with pro backing. Understanding all this stuff puts you in a better position to decide it you actually want to try.
With toy projects you can just keep iterating and iterating and never get anywhere, and not really fully realized why it's not successful.
I love Crafting Interpreters in part because he's explicit about this being mostly about learning and about play. I love languages (spoken, written, programming) and it was a great introduction that really got me rolling.
Articles like this consistently overwhelmed me, with the result being that I didn't really try it until years after I wanted to. Yeah, languages aren't for everyone, but specifically trying to scare people off of it seems harsh and unnecessary.
That depends. It's really hard to articulate just how much work is involved in writing a programming language. It's one of the classic infinite time sinks that exists in the CS field; no matter how much work you put into it, there's more to be done. It's very hard to ever call it "finished". I've seen a lot of language projects start as "I was working on this other thing and then I thought to myself, gee, here's a great opportunity for a programming language to make this easier." 8 years later they're still working on that programming language, the original project a distant memory.
Programming languages have a way of sucking in developers without them realizing it until it's too late. So while I wouldn't dissuade anyone from writing a PL per se, I would warn them that writing a PL is a *huge* undertaking in its own right, and should really never be undertaken to help you solve your own problems. If you want to go this route, then you need to abandon your problem and focus on the PL instead.
Personally I think an average programmer could knock out a BASIC, FORTH, or similar programming language - with extensions - in a few months of part-time work, if not less.
The harder part, the part that I can see taking a lot of time, is trying to create a "batteries included" language which has a very complete and complex standard library.
Sure there are cheats if you're able to use dlopen, and FFI, to access C-stuff, but it's still not easy to create client-libraries for MySQL, Postgres, Redis, HTTP, etc, etc. Those kind of extensions/support make your language very very useful for users though. So there's a trade-off.
I've put together a couple of simple languages, with a standard-library of maybe 30 functions. I'd draw the line at anything more no matter how appealing it might seem because I can just imagine how much of a time-sink it might become.
Lots of people could make a BASIC or a FORTH in a week. It would almost certainly be rather useless, unless the language itself is the point of the project using it.
99% of GitHub repos appear to be useless. They show up because someone wanted to make a simpler version of something that they can understand.
They don't teach much about the mainstream professional way. They often don't perform better, since things like GPU and SIMD are more important than simplicity.
They are perfectly fine as toy projects. I have a few myself(Although none of them use any real interesting algorithms or new concepts or anything). But that's... all they are.
At most they will become super niche things like suckless.
But all programmers should know that Electron apps run just fine even on cheap modern hardware, and that existing solutions are probably going to be better than anything you can do, unless you spend months to years like they did.
The programming community doesn't really seem to respect just doing everything by the book, the way a Microsoft dev would, and it's cool to see someone reminding people that it's a perfect good option to just grab some npm packages and make your thing and not bog yourself down with how it all works.
That is a bad thing. Going deep in any domain is hard; the thing to do is make the on-ramp gentler so more people can try it out, not make it steeper to discourage new learners.
If someone wants to start making games, we tell them to start off by cloning Pong or Breakout. We don’t say “be careful, here’s the long list of topics you’ll need to understand before you can make Fortnite.”
https://mukulrathi.com/create-your-own-programming-language/...
Also, the language descriptions are too brief to be helpful (Java and C# for being enterprisey, C++ and Rust for pointers and other system constructs)
https://en.m.wikipedia.org/wiki/History_of_programming_langu...
Today, however, Moore's law has stalled, and a different kind of power law is on the rise: core counts. Languages which previously saw success by being imperative and optimized for single core execution are falling over trying to keep up. We see increased efforts to harness the power of multiprocessing through asynchrony and parallelism being first-class citizens in newer languages, while others struggle to stay relevant.
Languages of the future will be built from the ground up to harness massive CPU core counts that will be available on consumer desktop machines in the near future. My advice to any budding language designer is stop looking to reinvent C and C++. There's currently a wave of programming languages that came of age circa 2010 - 2020 which are focused on just that (Go/Rust/Zig etc.), but if you're just looking to get into that space now you're late to the party. Instead look back to Erlang and the original promise of Object Oriented programming as the basis for a new language in 2022.
I’d bet future many-machine heterogeneous-resource languages will make that a lot easier.
To see the nitty gritty line-by-line walkthrough of everything that goes into actually building a language (all the way down to writing your own VM) I HIGHLY recommend reading Crafting Interpreters[1] by Bob Nystrom. I’m not a language hacker but found everything about this book worthwhile and very interesting.
The distinguishing features of it were the data model (it had first class access to the temporal data model we had on the back end), and that we kept the tri-state model of SQL with values and NULL.
Other than that it didn't have any real structured types, and just few high level actions.
It was a straightforward project, ALGOL-esque, basic control structures. It compiled to Java source code, and the resulting classes loaded in to the app server.
We used a "compiler compiler" tool, rather than doing it ourselves. It gave us an AST that we then walked. First time doing anything like that, especially with a tool like that.
Especially in an ALGOL like language, getting the expressions to work is the "hard" part. That's where your "napkin to white board to syntax file" falls apart with reduce errors and what not.
But when you get it to work, it's magic. When you get "a = a + 1" to work, and know that where 1 line of code works, 10000 lines of code will work, it's an amazing feeling. You just build the things a piece at a time, testing all the way.
In the end, we had an unresolved precedence issue (as I recall), but we never bothered to fix.
The funniest unexpected outcome was the classic SQL model of using NULL. Simply, 1 + NULL = NULL, and we expanded that to all expressions.
It kind of fell apart in something like: IF A = 1 AND B = 2 OR C > 3 THEN...
If any of those variable (A, B, C) was NULL, the entire expression evaluated to NULL, which was false. It took me by surprise when it happened, but, "duh", of course.
I simply changed the NULL rule to no apply to boolean operators. Instead of evaluating to NULL, I had them all evaluate to FALSE, and that fixed that.
Looking at Wirths work (Pascal, Oberon, etc.), you'll see that compilers can be simple. They're work, but they're simple work. We obviously make them more and more sophisticated all the time, but your compilers don't need to be that way to be effective, productive, and useful. We created thousands of lines of code using that language, it solved the problem very nicely that it was supposed to do.
Runtimes can be hard, but that's a separate problem.
> It kind of fell apart in something like: IF A = 1 AND B = 2 OR C > 3 THEN...
That is similar to how NaN values are propagated in IEEE 754 floating point standard.
However, a comparison with a NaN would cause an exception (in CPU or programming language, unless you're using a language such as C where FP exceptions are disabled by default).
And then there are several min and max operators that don't cause an exception on every NaN but which have different preferences for propagating NaN or values. The distinction is that a comparison affects control flow whereas min/max are still considered data flow.
That's not quite right. In SQL NULL represents an unknown value, so 1 + unknown => unknown.
TRUE OR NULL => TRUE
FALSE OR NULL => NULL
FALSE AND NULL => FALSE
TRUE AND NULL => NULL
Which is also why NULL = NULL => NULL, i.e. unknown == unknown => unknown.This is it right here. Ever since I started dipping into this stuff, it's been one of the most intoxicating drinks for me in all of programming. You just get a few little pieces working, and then you know you've created infinite permutations of working programs. Each feature you implement is another infinity of possibilities. There's nothing quite like it.
A minor typo spotted:
> it is impossible [not] to keep thinking about ways to improve your project.
Indeed the most productive days are those when I decide a major feature is unnecessary
And your comment resonates strongly with me. I have experienced exactly the same thing.
There is a time for 'the right' approach, and there is a time for crazy innovations that can take something to the next level.
What's their definition of "enterprisey"? It feels somewhat like an implied perjorative.
Maybe this could be rephrased as "study Java/C# to understand why many businesses choose these languages to build things that make them money"?
https://medium.com/hackernoon/considerations-for-programming...
Thanks for pointing out the riposte, I did not know it existed.
Amusingly, C++ has adopted a number of D's innovations!
I mention why because different whys have different orderings of hows. (For example, if you're interested in some new semantics, working on a parser first is probably counter-productive.)
Transpilers to C seem rare, maybe for a reason..
That said C transpilation is certainly a viable root (e.g. Idris does this)
But do also look into other parts of the LLVM frameworks, especially MLIR.
Historically, there have been many compilers that produced C code, but some may have done so because of lack of compiler frameworks. The first C++ compilers compiled to C. Eiffel compiled to C. Nim compiles to C. I've also used a compiler framework that compiled to Java.
It is interesting to consider the trade-offs even if you go with C or LLVM: