Crafting Interpreters: Superclasses
craftinginterpreters.com
craftinginterpreters.com
Adding a special VM instruction for such operations as inheriting from a class is a strikingly bad design decision.
You want to minimize the proliferation of VM instructions as much as possible.
A rule of thumb is: is it unreasonable to compile into a function call? Or else, is it likely to be heavily used in inner loops, requiring the fastest possible dispatch? If no, don't add an instruction for it. Compile it into an ordinary call of a run-time support function.
You're not going to inherit from a class hundreds of millions of times in some hot loop where the application spends 80% of its time; and if someone contrives such a thing, they don't necessarily deserve the language level support for cutting their run-time down.
> You want to minimize the proliferation of VM instructions as much as possible.
I guess that seems like half the trade off to me, or else VMs would all be OISCs. What you really want to do is to approximate Huffman encoding in your ISA, balancing that with ease of parsing (so bytes generally rather than bits).
With this VM, we have plenty of opcode space left, so it's natural to just add another instruction for inheritance, even if it does mean that the bytecode dispatch loop doesn't fit in the instruction cache quite as well.
I'll think about this some more. I'm very hesitant to make sweeping changes to the chapter (I want nothing more in life than for the book to be done), but this would be a good place to teach readers this fundamental technique of compiling language constructs to runtime function calls.
Maybe you could add a design note in the chapter with the mentioned content?
Keep it up!
Great idea!
Made a note of it: https://github.com/munificent/craftinginterpreters/issues/62...
That can be as basic as generating a function call AST node directly from a construct, or doing a node-to-node transformation.
The compiler doesn't see anything but a function call node (though any source code tracking info and whatnot still references the original construct).
If you're doing things with the AST other than just generating code, you may want to have a node for that original construct. (For instance, feeding cross-referencing info to a language server or whatever.)
But you're also using bytecode, which involves more caring about speed, so you've got a bit of a mismatch - as your code tries to match the multiple layers you're skipping, your code will get more complicated or your VM gets more arbitrary, so of the situation you have here.
This is the same approach taken by Lua and many of the original Pascal compilers, which had to run on very slow hardware, so there's pretty good precedent that you can get adequate performance.
> But you're also using bytecode, which involves more caring about speed, so you've got a bit of a mismatch
I wouldn't describe it as a mismatch as much as it is a trade-off. Because the compiler only has a peephole view into the source when it needs to generate code, many optimizations are off the table. However, because we have full control over the instruction set, we can sometimes tweak the bytecode format in order to more naturally align with the compiler.
In return for not needing an intermediate representation or AST, we get a much simpler compiler, especially in C where you have to worry about memory management.
Optimization along any axis (code size, simplicity, compile-time performance, runtime performance, memory size, etc.) always involves some level of violating software engineering norms.
One way to think of software engineering is that it's the practice of optimizing long-term developer velocity. Any practice that optimizes other factors generally does so at the expense of that. That's OK. Meta-software engineering is about knowing when long-term maintainability is the right factor to optimize for. :)
Technically, it's perfectly fine for a VM to be designed so that the semantic gap is small between it and the intended language family.
Many languages didn't support "arrow functions" AKA lambdas until fairly recently. Seeing how one would go about making large changes to the language is a good lesson to learn. One could argue that we should have known superclasses were going to be supported from the beginning, and I'd agree, hindsight is always 20:20.
I am struggling how this can be seen as a runtime instruction anyway - surely inheritance is a compile time semantic?
The new keyword requires a concrete class to instantiate. Abstractions such as interfaces cannot be used with new. The resulting instruction will directly reference that concrete class which makes it part of the software's binary interface. Its behavior can't be overridden: a constructor will always return an instance of the constructed class, never a subtype.
The factory design pattern was invented to solve these problems. Replacing new with normal method calls pretty much takes care of everything. The author himself has written about this on his blog:
http://journal.stuffwithstuff.com/2010/09/18/futureproofing-...
http://journal.stuffwithstuff.com/2010/12/14/the-trouble-wit...
I think a few explanatory notes and discussion of suitable alternative approaches is warranted but not a complete rewrite because its somehow wrong.
Each VM has its own quirks which produce their own opportunities. Why would this one be any different? Especially as a teaching instrument where it has to be simpler / clearer than a production codebase?
This is more difficult in a language which allows methods to be added to classes dynamically. IIRC both Python and JavaScript are in this category, but still calculate `super` lexically. Therefore, if a method is copied from one class to another, `super` within that method may refer to the wrong class.
Or does it only come up when a method is added to the class?
Possibly - I’m not sure about the details of method lookup in Lox. Does `this` even work correctly when a method is assigned to an instance?
If you take a function (which may be a reference to a method) and store it in a field on some instance, the function remembers its original bindings to "this" and "super", which is generally what you want and expect.
Lox doesn't allow you to mutate a class by adding new methods to it.
No, super() checks that "self" is an instance of the given class.
So if you take a method defined in a class hierarchy "A -> B", and attach it to an instance of "C" (i.e. c.foo = B.foo) and call it like "c.foo(c)", which you have to do, since you didn't invoke the descriptor protocol to bind the method, you will get a type error, because c is not an instance of A/B. That is also true if you attach the method to the class (C.foo = B.foo); you can call it like a method, but super() still checks the type and will complain.
Also, if you take a bound method and attach it to another object (i.e. c.foo = b.foo), it is already bound and you can't pass a "self" different from the object it's bound to. E.g. if you are passing around a callback that's a method.
Note how Class.method is just a function, while instance.method is a bound method. That's due to the descriptor protocol of functions (__get__/__set__).
You're right. Here's some example Python code:
>>> class A:
... def foo(self):
... print('A.foo')
>>> class B(A):
... def foo(self):
... super().foo()
... print('B.foo')
>>> b = B()
>>> b.foo()
A.foo
B.foo
>>> class C:
... def foo(self):
... print('C.foo')
>>> class D(C):
... def foo(self):
... super().foo()
... print('D.foo')
>>> d = D()
>>> d.foo()
C.foo
D.foo
Now, if I do `D.foo = B.foo` and call `d.foo()`, I do indeed get a "TypeError: super(type, obj): obj must be an instance or subtype of type". That's better than getting `A.foo B.foo`, as happens in JavaScript.However, if `super` were truly dynamic, I would have expected to get `C.foo B.foo`.
I think that would need an entirely differently designed mechanism and can't be bolted on.
edit: i should say i'm wondering which is more appropriate for me. my experience is that i took a PL class during my MS where we implemented a small recursive descent parser for a language (and the concomitant logic for evaluating the language). i'm interested in getting better at writing languages.
So if you're into the topic I think you can get something out of my books ([1], [2]) even if you've read Bob's book before. And if you read mine, then Bob's book will also show you something new.
I'm currently working through Bob's book myself (chapter 23) and I'm enjoying it immensely: the language, Lox, shares a lot of things with the one in my books (Monkey), like first-class functions and closures, but also has classes which Monkey does not. So now I can read Bob's book and on one hand think "Ohh, I wonder how he does _that_" and on the other hand "classes! let's see how that works."
My books also use Go exclusively and Bob's uses Java and C. I enjoyed using IntelliJ and writing Java for the first time in my life a lot and was always a fan of C, so that was also really interesting, to see how the ideas translate across three languages.
[0]: https://lobste.rs/s/43h7rz/crafting_interpreters_handbook_fo... [1]: https://interpreterbook.com [2]: https://compilerbook.com