I'm curious, how does that happen? Were the techniques used in compilers and interpreters of yore never recorded for posterity in academic papers or technical documentation?
The best stuff was all proprietary for decades. Even now, we only got Miranda's source code weeks ago. Miranda was written in the 1980s! It's the same for a lot of languages. Nial's the example I usually bring up for one that was freed far too late. People with the domain knowledge that would have helped deal with the complexity of modern systems are all either retired, have ditched PL-design, or have died.
Turbo Pascal, too, although that one's kind of strange because the author of it later went on to do a complete 180° turn while still working in PL design and implementation (plus it was never freed).
Ken Thompson's cc suite for Plan 9 was behind a million-dollar license fee until a decade ago, and it wasn't even for any modern CPUs. He once quit the industry to become a flight instructor; he's at Google, but he's retired now.
The best advice is "Stay small!" which most Free Software compilers seem to be violently against. It's not their fault, though: it's hard not to take new code when it's offered to you!
We'll probably never see the source code behind any version of the k interpreter by Arthur Whitney, and trying to get an interview with him is like going on a snipe hunt, so good luck figuring out anything about his approach besides "Small!" It doesn't help that Kx sues the hell out of anyone who's actually seen the k source and tries implementing anything remotely like it that's not a toy.
Every modern compiler is for multiple architectures in convoluted ways. This is generally a terrible idea, especially with how divergent they're getting in the times we're in, even between chips that have the same instruction set.
Intel x86_64 processors aggressively speculate almost as intensely as Transmeta did back in the day, but we're still treating them and AMD chips more or less the same, and we're using the same compilers that we're using on that instruction set for RISC-V and for ARM and for Itanium and for obscure 16-bit CPUs and so on.
Portable compilers aren't bad in principle, but when you look at pcc compared to GCC you'll see exactly where we went wrong: pcc wasn't very optimized, it was simple, and it was understandable. Great for bootstrapping. Not great for getting the most out of your CPU. Having portable compilers that try to heavily optimize is a mistake.
No one seems to know what a CPU cache is!
And how many compilers do you know that compile to C, or to LLVM bytecode or similar? It's insane: nobody should be doing that! That approach doesn't make much sense!
And the only argument any of these people can make is "But our language is too big for it to be practical to write something unique for each architecture! Piggybacking makes it way quicker!"
No one seems to realize that their languages are getting too large. They aren't even getting too large in a graceful way: Common Lisp is probably the biggest language there is, but it's extremely portable, and is easy to write an interpreter for.
Things got complicated really fast, and domain knowledge was sort of lost in the waves of proprietary compilers as they were destroyed by GCC. I appreciate that free software "won" to some extent, but if it hadn't have won as fast, we might have had some knowledge transfer happen that was actually useful.
Walter Bright isn't doing magic, but he and the people he work with seem to be some of the only people in Free Software who are making a compiler that's fast and good in the x86_64 world.
It's also the approach used by Open64, where you compile to WHIRL which is then optimised with successive passes until you eventually output whatever the architecture can run.
BCPL's O-Code is far different from LLVM bytecode. O-Code was little more than a virtual stack machine. Very similar to Forth, though slightly more complicated.
I could be wrong, but isn't Open64 more or less dead? The only thing I can think of that actually uses it (CUDA) isn't best-in-class.
Xerox PARC had Smalltalk, Mesa (later Mesa/Cedar) and Interlisp-D.
The platforms would first load the respective microcode into the CPU for the bytecodes used by the desired workstation environment.
Initially Xerox PARC made use of BCPL, but quickly they realised it wasn't the best way to write systems software and created Mesa to replace it.
After all, BCPL was intended to bootstrap CPL, not to write fully systems with it.
Other 60's computer companies were making use of either Algol or PL/I dialects, all the way up to the 80's.
Lots of juice papers at Bitsavers.
1) why miranda was fast itself (was it compiling perf or runtime perf) ?
2) why haskell is so bad .. since miranda lineage with haskell is quite strong, unless SPJ, JH, PW and the likes had zero access to miranda techniques or no legal right to use them.. I fail to see how they made haskell compiler so bad.
If I’m being paid to write a Ruby compiler then ‘keep the language simple’ isn’t an option available to me, is it.
And the techniques from the 80s would be absolutely hopeless at compiling and optimising Ruby. A lot of the implementation approaches you’re talking about work brilliantly for a single-pass compiler for a trivial language like Pascal generating basic machine code but they aren’t expansive or powerful enough for bigger problems.
So the techniques haven’t been lost - they are in most cases either not enough or not applicable to our languages and the output we need.
"the art of writing compilers and interpreters has effectively been lost"
I'm curious, how does that happen? Were the techniques used in compilers and interpreters of yore never recorded for posterity in academic papers or technical documentation?
I think I did a pretty good job in answering it, although it was a bit rambling.
The people on the standards committees for these languages are generally also the ones implementing compilers, at least at first. The ones that don't have significant overlap between the two fail. ALGOL-68 comes to mind in the past, FORTRAN-03/08 comes to mind now.
Also, there were fast Lisp compilers back then (as fast as Lisp can get, at least, and keyword "compilers" as I'm not misinterpreting the two), along with decent (eh) Smalltalk compilers: Ruby isn't that unique. An optimizing compiler for Ruby could have used many of the techniques, given that most of the bottlenecks in Ruby are the same as many Lisps. Of course, it won't map perfectly, but much of it would. (I will admit that this paragraph is written with primarily old versions of Ruby in mind, because beyond 2007 or so, when it was basically a Lisp with Smalltalk tendencies under the hood, I haven't really kept up.)
Many techniques absolutely have been lost over time, even in obvious spots, where you'd expect them to be maintained quite well; I have a folder of poorly-OCR'd, poorly-scanned ACM papers of novel, technically-interesting & fast pre-1990 compilers with significantly less downloads than citations. Academic papers aren't really where you'd expect to see "lost" techniques, but publishing is a wasteland, and most non-famous on this subject (and computers in general, partially; ever tried to find more than a handful of papers on Sun's Spring?) written between 1970 and 1990 are practically lost. The industry has the memory of a dormouse, though it might get better now that more and more things are becoming source-available.
Well, maybe. But it is hard to believe if most of those things aren't rediscovered in these decades. And IIRC that GCC took over partially because it was quite better. GCC / LLVM appear to implement every new technique in someone's PhD thesis like Diophantine equation solvers into optimization passes. I suspect it is no longer the case.
If you're saying compiler itself is slow, well the rat race among compilers is producing benchmark binaries (sadly). But given how much we have advanced, I suppose fine grained incremental compilation should be norm in every compiler.
> Turbo Pascal, too
Here I suppose you are talking about compile speeds. Well Turbo Pascal had the advantage that it was integrated with an IDE. Eclipse and IIRC visual studio can also reach those speeds using incremental / background compilation.
> We'll probably never see the source code behind any version of the k interpreter by Arthur Whitney
I have heard a about this K. What is this? Is this an array programming language like APL? I would like some pointers.
> Having portable compilers that try to heavily optimize is a mistake. Again here there are tradeoffs. People like the comfort of portability and thus don't care about small performance increments as long as code performs reasonably well. Although I would think we can have a peephole optimizer that can optimize based on target machine upon installation.
> No one seems to know what a CPU cache is!
While this may be true for normal programmer, compiler implementers care quite a lot about caches. Especially most JIT optimizations are about cache.
> Compiling to C / LLVM IR doesn't make sense.
Sure that's a tradeoff. But that allows to implement compilers in less time and utilize most optimizations.
However I too think the progress in compilers has stalled in these years. There are native programming languages being created. But it is almost Compile to C or embracing LLVM monoculture. Appreciably Go didn't do that. And due to proliferation of scripting / JIT languages, there's not much scope for optimizing or fast compilers. In particular I am sceptical of LLVM monoculture. The compilers rat race prioritizing benchmarks above other things isn't going to end well. Moreover my gut feeling is today's compiler infrastructure is not suitable for High level languages.
I was recently listening to the 2007 "Copyleft Capitalism: GPLv3 & the Future of Software Innovation" talk¹ by Eben Moglen (who incidentally happened to have worked on IBM's APL interpreter), and he makes the opposite argument: it was the source code hoarding of the Microsoft-driven personal computer industry that was the biggest impediment to knowledge transfer and innovation.
But free compilers won so quickly even when they weren't obviously better that many proprietary compilers died quick, unceremonious deaths, with no obvious successors, and leading their authors to go into different areas. This was a bad thing, in my opinion, because a lot of early compilers were written in incredibly clever ways, and those clever ways died with them.
I think my comment was also worded poorly:
> I appreciate
> that free software "won" to some extent
is how I was hoping it would be interpreted, but I think it was interpreted as:
> I appreciate
> to some extent
> that free software "won"
Free software winning was absolutely the right thing, I just wish it would have been a bit less sudden.
Intel Fortran compiler and their C++ compiler produce the fastest code. Lucid Energize and IBM VisualAge C++ did incremental compilation and had IDE features not available with Free Software C++ compilers and editors/debuggers.
https://www.youtube.com/watch?v=pQQTScuApWk http://www.edm2.com/index.php/VisualAge_C++_4.0_Review
It's not exactly hidden knowledge that the Intel compilers produce the best code, it's just that they have draconian licensing requirements. Furthermore, that knowledge cannot possibly be lost, because they're still actively developed to this day. I'm interested in the technological specifics of what these older platforms did that the newer ones cannot do, and why the knowledge of how to do them has been lost.