Compiled and interpreted languages: Two ways of saying tomato
tratt.net
tratt.net
[0]: http://blog.sigfpe.com/2009/05/three-projections-of-doctor-f...
[1]: https://news.ycombinator.com/item?id=18995651 (and the parent: https://news.ycombinator.com/item?id=18995498)
1. You're interpreting if syntactic tokens directly drive evaluation. So only first two are interpreted.
2. Interpreters take syntactic input and evaluate to a result that comes from running the program, a compiler takes syntactic input and outputs a program that must then itself be run to compute the result.
That said, I agree that "interpreted language" or "compiled language" is not a formally correct term, it's a connotative definition implying what's typical or the purposes for which a language may be suited.
if you see the program's inner loops, it's been compiled
if you see the language's inner loops, it's being interpreted
ie, in the latter case, you'll see the language's dispatch loop in the "hot" range, as it interprets the def'n of qsort at multiple recursion levels... (and this signature will be independent of whether or not calls in each level are made on their own or combined horizontally to gain nested parallelism)
Consider a simpler example: let's profile not just qsort, but a couple of other relatively long-running library calls, and let's strip the executable beforehand, so we aren't even working with symbolic information, just a bunch of code addresses.
If we get multiple different hotspots, we're likely looking at a compiled library.
If we get the same hotspot in each case, we're likely looking at an interpreted library (and seeing the interpreter's dispatch loop).
One of the nice perspective insights from 3Lisp, which of course was very high level!
> Another thing I hope that I've indirectly shown, or at least hinted at, is that any language can be implemented in any style. Even if you want to think of CPython as an "interpreter" (a term that can be dangerous, as it occludes the separate internal compiler), there's PyPy, which dynamically compiles Python programs into machine code, and various systems which statically compile Python code into machine code (see e.g. mypyc).
(Emphasis mine:)
> There are interpreters for C and C++ (e.g. Cint or Cling).
Usually those are for (large) subsets of the language though.
https://github.com/graalvm/sulong/blob/master/docs/ARCHITECT...
"just add '#!/usr/local/bin/tcc -run' at the first line of your C source, and execute it directly from the command line."
It's all about transformation of code to other code... and we can be more flexible in our classification. If modern languages even need one.
lisp in small pieces does try to convey the idea that interpretation and compilation are siblings (by partially deriving a transformer from an interpreter and ultimately a bytecode vm) but it would be worth a full book.
However, JVM bytecode is actually both interpreted and JIT compiled depending on runtime performance and other parameters. The JVM can even switch between compilation and interpreting for the same piece of code multiple times (most commonly, when attaching a debugger and hitting a breakpoint).
This analysis focuses on the language specification and ignores the wider picture.
The first question we have to ask about a language is what caused it to be created. Making a language is a lot of effort (though it can be a lot of fun), and there has to be a reason for it (even if it just curiosity).
Now to simplify stuff, let us choose to ignore purely exploratory languages.
The first question to ask is what problem the author was trying to solve with the language.
Take a look at C and AWK in which Brian Kernigan was heavily involved. The design goals for C were a lot different than the design goals of AWK. This led to C generally being compiled, and AWK being interpreted.
In addition, in the case of C, there were deliberate decision decisions made to make it easier and to compile and optimize (for example all the undefined behavior).
Now let's take a look at Java. The solution space that this was aiming for from nearly the beginning was a high-performance language compiled to a virtual machine using garbage collection. A lot of design decisions were made for the language with that goal in mind.
Or look at C++. The design goals basically necessitated a compiler (even if the output of the compiler was C as CFront did) vs an interpreter.
In addition, an ecosystem of a language is much more than the language, and can even be more important than the language itself. And the ecosystem, for the most part picks a side in the interpreted vs compiled debate.
You could have a Java without the JVM, but you would lose access to a lot of techniques, libraries, debugging tools, and development tools that you have now.
C++ does have interpreters, but these are all "use at your own risk" and many libraries will behave in weird ways (with regard to static initialization, etc) since they are being used in an unintended manner. I don't think anybody uses an interpreted C++ in a production system, and rather it is used more for exploratory development and analysis.
Similarly a compiled python loses compatibility with certain libraries and there are issues that pop up that you would not have with the more widely used interpreted version.
So yes, theoretically languages can be compiled or interpreted, but practically from the design going forward, usually one or the other is explicitly or implicitly aimed for, and using a language against that grain can lead to a lot of unnecessary pain.
It then gives you tooling to poke around the bytecode which is clear enough that you can read it. I once gave a talk showing x64 and bytecode side by side for some arithmetic function and if you pick the example carefully they line up well.
Then there's cython and pypy to point to for different points in the design space than cpython, or jython if you want to talk about the jvm. There's probably one on .net and iirc truffle has an implementation as well. The abandoned unladen swallow. So there's loads of python implementations all approaching the interpreted vs compiled tradeoffs differently.
The difference is most obvious looking at arrays of fixed-size integers, where a C int* far outperforms a Python list of ints no matter the implementation. Python list is actually a different thing because it has variable-sized integers and can take other objects too, so you can't optimize around that but could do so with a more specific fixed-size int array interface... which is what NumPy did.
The author has identified that "compiled" and "interpreted" don't mean much about how the code technically runs. What really makes the difference is the intended use case. Same reason tomatoes aren't fruits in layman's terms.