Musings on AOT, JIT and Language Design
pointersgonewild.com
pointersgonewild.com
[1] http://pointersgonewild.com/2012/11/09/static-vs-dynamic-why...
> [in C] It's very difficult to avoid undefined behavior,
True, but LLVM IR and other IRs have undefined behavior as well.
> you won't be able to modify the compiler backend (and you will want to as you start optimizing),
Not all languages need backend changes. And you can still add IR optimizations or backend changes that help your language, if you do. Yes, this might not be as easy, but then if you compile to LLVM IR, you might need IR changes for your special things anyhow.
> you can't add new language intrinsics, good GC is pretty much incompatible with that,
Good point, for most GC languages compiling to C is not a good option, which makes sense since C has no GC support.
> and precise control over debug info is impossible.
Yes, but using the C preprocessor you can get pretty far.
> Compile to LLVM IR and/or GCC IR instead.
More power that way, sure, but
1. It limits you to one compiler, while C has many.
2. LLVM IR changes over time (and recently had plenty of examples of this), so you'll need to track that. Whereas C is extremely stable, so long-term, it's much less work.
3. Emitting C is very easy (for those familiar with C), and also very easy to debug.
I think compiling to C is an excellent option. Sometimes better, sometimes worse.
> True, but LLVM IR and other IRs have undefined behavior
> as well.
LLVM IR has less undefined behavior than C, and even despite the reduced surface area frontends have still historically wrestled with taming its UB (both PNaCL and Rust have documented their struggles). > It limits you to one compiler, while C has many.
To quote Walter Bright on this topic (which I've linked elsewhere in this thread, but here it is again: http://forum.dlang.org/post/n1vbos$11ov$1@digitalmars.com): > 1. You're at the mercy of bugs in the C compiler you
> cannot fix.
> 2. C leaves quite a lot as "implementation defined",
> causing endless compatibility issues with various C
> compilers.
> [...]
> 6. You'll suffer from endless bug reports caused by a
> mismatch between your compiler and the user's C
> compiler, whatever that might be.
> [...]
> 10. The order of evaluation of C code expressions is
> implementation defined.
> 11. Installation problems, again, because you don't
> control the user's C compiler.And there are also non-semantic differences that end up being important, like one browser running certain things too slowly or taking too much memory.
Multiple implementations is always more complicated ;) But in many cases it's better than a monoculture, in my opinion.
Meanwhile, the only thing that you gain from all this extra effort is the ability to target platforms that GCC and LLVM don't already support, but these platforms are all likely to be obscure resource-constrained devices, and I just can't see developers on these platforms flocking to such a thoroughly-dynamic language as the OP is envisioning.
But yes, I agree that in the context of a really dynamic language, it might not be what the author wants anyhow.
The very amount of tests in the LLVM source tree ensures that nobody will be motivated to do any significant change unless it is really, really needed, otherwise you'd be forced to do a lot of painful work to update all the tests.
Also, compiling to C pretty much rules out having a REPL for free and all the other fun, while with LLVM it's trivial.
* the removal of pointer types (still happening)
* the use/user flip
In my experience these have been significant in the amount of work it takes to track.
In fact, there isn't any dialect of C, standard or not, that I'm aware of that allows you to do this in a fashion that works with precise, moving GC. (Mostly-copying doesn't count.) Maybe C++/CLI, but the object model you're opting into with that brings you so far from C as to barely be considered "compiling to C".
The problem is not so much the spilling and reloading (though that is a large issue) as what that is going to do to the ability of the optimizer to reason about the code.
(If you're curious how I'm so sure it's a huge performance penalty, I have been spending a fair amount of time of the last month working with others to debug large performance problems caused by a variant of exactly this technique used in Servo.) :)
We have some macros and a bunch of functions then that make emitting LLVM IR simpler. For example, the "ins--iterate" and "ins--if" / "ins--else" macros in this code handle phi nodes and so on:
https://github.com/dylan-lang/opendylan/blob/05271f6fec9da05...
This library is currently in the process of being updated from LLVM 3.5 to current HEAD as debug info has changed significantly upstream. After that, we will be landing some documentation for it as well.
If you're interested in learning or hearing more about this, we're typically around on #dylan on freenode IRC.
The LLVM GC interface is controlled through intrinsics, so you can access it through the human-readable IR representation just as you can with the API.
The obvious one is not having a giant build time dependency that might change and be at different versions on different platforms. The text interface is far more stable.
The problem I had with doing this is how picky llvm is when it comes to ordering your temporary registers.
Obviously if you enjoy making languages then you should do so - and in that case, noticing that your dynamic un-toolable language won't get wide scale adoption shouldn't matter. It's already a guarantee it won't get that wide scale adoption. Either make your un-toolable language because you enjoy making them, or contribute your efforts to an already popular language that you enjoy. There are tons of interesting problems and huge areas for improvements in every single language.
Don't make a new language for the goal of gaining widescale adoption.
I suspect what you're trying to say is "don't expect to succeed." That's true--most people will fail, but we should all hope people continue trying, unless we're satisfied with what we have. I'm not satisfied, and I think progress will require new languages, not just improving the ones we have.
Some things require that a lot of people try and fail. Language design is one of them.
That's what I'm reacting to, and I still disagree. People should have an eye on what getting traction would require, even if only a tiny fraction of language designers can make it.
However even more importantly - if you're going to create a language, it needs to be the ultimate language in your own eyes. You can't leave off novel features that you have experience in because it doesn't have mass appeal or has tooling issues. You need to create the best language you possibly can if you want to compete with others trying to create the best language. Hamstringing your enjoyment of the language by going in a direction you don't like because it has more apparently mass market appeal isn't any more likely to be successful, and it certainly isn't going to be enjoyable enough for you to slog through the tiny complexities.
While there is a tiny chance that your language becomes the next big thing, it probably won't. That means you need to make the language you want to make and hope it succeeds, not make the language that 'will succeed' but you don't want to make.
I agree with you so far as it can't be a language you don't want to make. I don't think this article describes that, but a more modest set of constraints. Elsewhere, the author mentions Julia as a positive inspiration.
I think the idea of compromise depends on the idea you're trying to bring into the world. You can't compromise your core idea, but you can compromise on other things. I guess I'm making assumptions about what the author is compromising that I'm not entitled to. Perhaps you know that I'm wrong.
We need people that look at the experience of programming as a whole, writing, editing, debugging, profiling, testing, analysis and packaging. People that take all of that into consideration and design languages within that context and under the constraints that imposes.
If you only see yourself as a language designer and ignore the wider context, your language will fail.
To be sure, they do program applications that are quite different from what 'bread-and-butter' programmers do, and that might influence their taste in programming language features.
We need languages designed to be used by future generations of programmers in a world where usability and tools come first. The "Manipulate symbols in this arcane syntax and hit F5/Run until it does what you want" experience is fundamentally at fault. Sure we have debuggers and IDEs, but they're more or less tacked on top of core flaws, things that allow for whitespace formatting debates or semicolon religion. My IDE and my VCS should just edit and encode the AST of my language. The compiler and IDE should share the exact same code for indexing and finding symbols in that AST -- the language acting as a library to the IDE so new releases with new features are automatically understood instead of lots of duplicated language for highlighting/completing/searching every language for every IDE. These are just a couple examples of the ways we're currently living in a state of sin.
The standard library of a language should provide access to the parser, so that developers can easily create tools for the language they're using.
1) It's not "trivial" -- it's actually a substantial amount of work, especially as the language evolves. How many parsers have you written? Do they have public APIs? Which public APIs to parsers have you used before?
2) It can actually constrain the evolution of the language. IIRC Python has had some non-trivial upgrades that "reinterpret" text with a different parse. If you have to provide parse tree compatibility in addition to text compatibility, that makes changing things harder.
I've noticed that people who have taken a compiler class (often not in C) but haven't worked on "real world" parsers (nearly all written in C/C++) have a very idealized notion of how they behave. Parsers written in C are more often than not tightly coupled to both the lexer and compiler.
You can turn the argument around. As writing a parser is near trivial with parser generator tools, parser APIs are not what is problematic with writing tools.
People have been saying this for generations. Nobody seems to really believes it enough to step up and try to make it work.
Manipulate symbols in this arcane syntax
People have been experimenting with alternatives to formal languages for decades. Some of the results might have been helpful for beginners, but as tool for expert programmers, formal syntax might not yet have been beaten.The same goes for language designers. I don't think that they have no experience programming, they obviously do. I just don't think that it matters all that much.
Additionally the experience of developing software varies widely. Some use IDEs, some don't, some people use debuggers, others use print statements, some people use profilers, others never have a need to. Even if you have programming experience, your particular experience will be very different from the experience other people have, even developing in the same language.
So no matter how little or much programming experience you have, you can't create a great programming experience based on your experience alone. You need to have an understanding of other programming experiences.
You need to have an understanding of other programming experiences.
Part of the problem of designing programming languages is that there is little agreement and extreme heterogeneity in the opinions of what is and what isn't a good idea in programming language design. Ideally, PL design statements like "language X is better than Y" should be accompanied by rigorous empiricism. Clearly, the field of PL design has a measurement problem. PL design doesn't have empirically validated and methodologically sound studies that quantify the "dividends" of one PL or PL feature "over another, such as reductions in defects, program size, or development time.Of course PL development isn't the only field that lacks easy empirics. And PL development shouldn't let the perfect be the enemy of the good. It's quite clear that PL features do have positive productivity consequences, with garbage collection being the most clear-cut example.
The interesting question is what do we do the absence of good empirical methods.
There are several valid answers to this question. One of them is to be strongly guided by mathematical beauty and elegance. This is the family of languages that started with Lisp and is today represented most purely by Haskell, Scala, Clojure, Ocaml and F#. The key ideas from this family: garbage collection, higher-order functions, rich types, have been accepted by the programming mainstream.
One wonders if these ideas would have been discovered if John McCarthy had studied the programming experience of working Fortran / Cobol programmers, rather than Church's purely theoretical lambda-calculus.
Further than that, in language design you obviously can't do A/B testing, since the cost would be too high, focus groups are in general fundamentally flawed because you can validate whatever you want and UX studies in language design would be too superficial, unless performed for a long period of time, but then you get to the same problems the scientists have to deal with in studies on nutrition.
So unless you're proposing design by committee, which never worked in language design, you tend to get the best results from one language designer that outlays the foundation based on personal opinions and ideally on the prior work of others and is then helped by the community.
I do understand your point on different styles of programming. However you can't please everybody. For example if you choose for your language to be IDE friendly, you end up with Java. Not sure how many people remember, but one reason for why Java EE has been a complete clusterfuck was the IDE oriented design, as in this belief that an IDE can help with the boilerplate and the wiring. The availability of IDEs is also why XML become popular. Even today Java is not usable without an IDE, because Java's libraries have been designed with IDEs in mind and they are too verbose for your brain to even remember the names of methods, let alone the happy path. Ruby's libraries on the other hand have been designed for expressiveness, for being easy to remember, also to counteract the poor IDE support. Neither approach is superior, but the choice in design was very conscious.
Honestly, I'd reverse it and instead of looking for exceptions, ask which languages the OP was thinking of? I'm sure they exist.
Most of what we're doing isn't about what's technically the best so much as what was and is popular. That's for backward compatibility, integration, network, and marketing benefits. A clean slate attempt at profiling-guided AOT compilation, at server or binary, could probably match or do a lot better than most modern stuff.
I would love to see a language that breaks this mold. Where the language and the tools are intimately coupled, or at least had a more complex API. Why aren't the concepts of the convention over configuration movement being pushed further, into the developer workflow and the language itself.
What if unit tests exercising your code was expected, because this is what drives code completion.
How about a community that won't accept a new library as being complete until it hooks into the standardized tools to support its custom macros.
I'd love to see what could happen if we question our basic assumptions of the roles between language, tool, and programmer. What could we do differently if a program was stored as pure AST, instead of a text file, for instance?
Dynamic, run-time reflection - yes, it is very harmful indeed. But dynamically extensible parser per se is not an obstacle for tooling.
> If you can hijack the parsing of a program while it’s being parsed, and alter the parser using custom code you wrote, then good luck running any kind of static analysis. It’s just not going to work.
It works. Easily. No problem at all. As long as you can clearly separate compilation time and runtime.
I could easily make any dynamically introduced syntax extension to play well with an IDE without any user effort. Autocompletion, semantic highlighting, type balloons, autoindentation - all the usual IDE stuff.
> Rest assured though, this language will be dynamically typed ;)
This is a big mistake. Dynamic typing breaks tooling, not an extensible syntax.
And Lisp is far less dynamic than Smalltalk and the modern lesser toys like Python.
See in ACL2 how a fully static typing can be added on top without a significant change in semantics.
If I were doing a new language I would compile to web assembly.