Stage polymorphism I guess means abstracting over how many stages there are till you get to plain old non-code data; hopefully someone will chime in who's more up to date or free to read the paper tonight.
eval: "expression of type U" -> U
quote: U -> "expression of type U"
chriswarbo's reply higher up should be taken to supersede my reply -- he's obviously more current on this stuff.No, it's a type. 'list' is a type constructor. Staging is a different beast from types altogether. Read up on MetaOCaml for how staging works in a typed language.
You're probably confused by the fact that typed languages assign types to expressions, but a value of type "expression" is something different. You're reifying the AST of an expression as a value at runtime, and then you can build further expressions, and then compile them all.
A type is not an expression. We wouldn't have two words designating the same concept. Even in dependently typed languages where types and expressions are intermingled, they are still distinct concepts.
Now, you can sort of talk about expressions in the "type language", but these are not expressions of the "value language". Even so, an unqualified statement like "a type is an expression" is simply incorrect because "expression" always refers to the value language, so that phrase conveys the completely wrong intention.
Finally, as for how to relate staging to concepts you might be more familiar with, I suggest the paper, Closing the Stage: From Staged Code to Typed Closures [1].
This is only true in dependently typed languages. And even then, soundness requires stratifying types into universes or something similar.
> In other words, types are higher order expressions.
This isn't the meaning of higher order as it applies to programming languages or the standard isomorphism to logic.
From reading section 3, it seems that "stage polymorphism" allows the same piece of code to be used in different "stages". For example, we might have a function call like `square(4)`: if we evaluate it now, like an interpreter, we get the value `16`; if instead we "stage" it, like a compiler, we get code which (when executed) will call `square(4)`.
The polymorphism comes from parameterising the 'elimination forms' (branching, function calls, etc.). We can think of `square(4)` as being `call(square, 4)`, and we're overloading the choice of `call`: for an interpreter, we use a `call` which does the function call now; for a compiler, we use a `call` which constructs code for doing the call.
As for regular expressions in Javascript, this is more powerful for several reasons. Firstly, regular expressions are so limited that they can't reference other values; hence there's not much difference between interpreting or compiling them.
What about a more powerful embedded language, like `eval` running Javascript from within Javascript? That has the problem that we can't send values between different "levels" of Javascript. Say we have a value `x = 42` and we want to create an 'embedded' program `x + x`. We can pass around a string `"x + x"`, but when it eventually gets sent to `eval` it won't necessarily use the same `x` as we intended (it basically suffers from dynamic scope).
If we had a way to "stage" Javascript from within Javascript, we could ensure the correct value is used, but we'd probably have to write some funky expression like `<,x + ,x>` (depending on the language; take a look at MetaML for an example!). If we want to stage some Javascript which stages some Javascript (and so on), we'd accumulate horrible nesting/escaping boilerplate.
This "stage polymorphism" lets us write `x + x` for all stages, including things which are evaluated immediately. Their technique is also one pass, meaning that we don't have to run evaluators in compilers in evaluators... It also works with reflection, and with interpreters which implement the language semantics differently (they include examples like maintaining a count of how many times a variable is accessed, and for converting to continuation passing style).
If we think of a classic OOP example, we might say (in some made-up language):
Mammals can breathe
Mammals can move
Dogs are Mammals
Dolphins are Mammals
From this, we know that Dogs and Dolphins can breathe and move, so we can write code like: function checkStatus(Mammal m) {
try {
m.move();
return "Free";
} catch {
try {
m.breathe();
return "Trapped";
} catch {
return "Dead";
}
}
}
This code is subclass polymorphic, since we can pass in a Dog or a Dolphin, or some other type of Mammal, and it will work unchanged.Yet this is completely orthogonal to inheritance! There are two ways we might actually implement these constructs:
- Mammal is an interface: Dog implements breathe by operating its lungs and move by operating its legs; Dolphin implements breathe by operating its lungs and move by operating its tail and fins.
- Mammal is an abstract class which implements breathe by operating its lungs. Dog inherits breathe and implements move by operating its legs; Dolphin inherits breathe and implements move by operating its tail and fins.
Both of these are valid approaches, but consider that:
- Only the second approach uses inheritance, so it seems strange to limit "polymorphism" to only this case.
- The checkStatus code doesn't actually care which approach we take. It's "polymorphic in our choice of polymorphism"!
Also, let's say that we did restrict the term "polymorphism" to these sorts of mechanisms. We might say that the above example is "polymorphic in the choice of Mammal". In which case, the "stage polymorphism" of this article is nothing other than "polymorphic in the choice of stage".
We could implement it in a language like C++ something like:
Stages can call functions with arguments
Stages can branch on booleans
...
Interpreter is a Stage
Compiler is a Stage
We achieve polymophism by passing around the currentStage, and writing code like `currentStage.call(myFunc, myArg)`. If `currentStage` is an `Interpreter`, the call will be performed immediately. If `currentStage` is a `Compiler`, code will be generated to perform the call.The only difference between such an OO setup and the actual implementation in the article is that boilerplate like `currentStage` is all handled implicitly, rather than manually passed around, and we overload the normal language constructs rather than replacing them with methods, e.g. we can write `double(foo)` instead of `currentStage.call(double, foo)`, and `if foo then bar else baz` instead of `currentStage.branch(foo, bar, baz)` (or something even worse, to ensure that `bar` and `baz` get delayed!)
In C++, an image of the data members of a base class is embedded verbatim inside a derived object. This means that a function compiled to work with a pointer to a block of memory formatted as a Base object can just as well be passed a pointer to the region of memory corresponding to a Base object inside the larger block of memory allocated to a Derived object. So C++ saying that inheritance implies polymorphism is not because they are conceptually the same, but rather that an implementation exists that can give both features at once for little runtime overhead.
Also, C++ doesn't distinguish between interfaces and classes; an interface in C++ is just a class with no data members.
Runtime polymorphism in C++ requires marking methods as virtual. They are not virtual by default. Why not? Because the default assumption is that code can operate directly (and efficiently) on what it statically knows about an object, and in the case when none of the methods are virtual it allows the object and code to be slimmed down slightly.
C++ has many other forms of polymorphism, method overloading or operator overloading (a.k.a. ad hoc polymorphism), templates, etc. The reason they are not all treated as the same is because their implementations are different. Hiding all of the differences would impose a small but fixed cost on all of them, which would be counter to C++'s principle of you don't pay for what you don't use.
> Also, C++ doesn't distinguish between interfaces and classes; an interface in C++ is just a class with no data members.
This doesn't invalidate the point about polymorphism being orthogonal to inheritance. The point is that we can choose whether or not to use inheritance: `breathe` could be implemented in `Mammal` and get inherited by `Dog` and `Dolphin`, or Dog and Dolphin could implement `breathe` themselves and there would be no inheritance. Either way, we still have (subclass) polymorphism, so saying "polymorphism should mean inheritance" seems like a bad idea.
A solution is to make the function which builds the expression object be (1) a callback and (2) polymorphic in the staged sense as you described, where the exact same code 'x + x' can have different meanings in different stages. In the first pass, the objects the callback works with have operations which are defined to just build an expression object without evaluating it. In subsequent passes, the same code (or a differently parameterized specialization of the same generic code) is called again but this time it is manipulating objects bound to actual runtime state. The callback may or may not be actually doing any computation (perhaps the interpreter can or is evaluating the expressions itself, and the operations the callback is doing are all nops). But just the act of re-threading execution back through the callback function gets the expression builder function's name back onto the stack for exception stack traces. It also provides a sort of "high level user interface for debugging": if the user puts a breakpoint in the callback code, examining the variables in it gives information about the runtime values passing through the interpreter. I.e. it would provide a way for the user to directly look for temporary variable "x" in the interpreter's current running state.
I guess that's in some sense the opposite of collapsing stages? In a way of thinking this is putting stages back in after they have been collapsed, for the purpose of tracing execution.