Aside: That was a very thorough and well written blog post, mraleph :) A very enjoyable read.
edit - I do think the blog post implies that v8 can't do polymorphic inlining. That's the impression I got from it anyway (and I was somewhat surprised, as I had assumed it would). It may be worthwhile to change the wording to make it more clear that V8 can inline more than one callee at a site.
It took me a bit to realize where this confusion comes from (especially given the whole discussion of how polymorphic property access is handled). Is it due to "Undiscussed - Not all caches are the same" section?
Indeed V8 can't at the moment do polymorphic inlining at non-method callsites - f(), however it can on method callsites o.m().
The reason for this is lack of type feedback.
I will seek to clarify the wording in that section.
> Aside: That was a very thorough and well written blog post, mraleph :) A very enjoyable read.
Thanks.
Not sure if it's worth the effort, though.
I think most of the reason SM has polymorphic inlining for all callsites is that we added it first as a stop-gap for methodcall inlining, and then later went back and added support for method callsites (guarding directly on the receiver object's type and eliding the property lookup - which I assume is what V8 did from the start).
If you already have methodcall polymorphic inlining, the return on investment in expanding that to all callsites is perhaps pretty low.
Would be interesting to measure though..
There is certain technical heritage here.
Originally method calls compiled down to a single IC that did load and call within a IC stub. Type feedback from these ICs was interpreted in the same way as from property load ICs. On the other hand function calls cb(...) compiled down to a call through CallFunctionStub which did not record any feedback whatsoever. At some point CallFunctionStub learned to record a bit of type feedback (monomorphic / megamorphic) and we started using it in Crankshaft for speculative inlining.
Only recently method calls were decomposed into Load IC and a separate Call IC (an evolved CallFunctionStub) which invokes loaded function. For property invocation o.m(...) feedback comes from the Load IC now, and Call IC feedback is ignored. For variable invocation cb(...) feedback comes from the Call IC - which is still only able to distinguish only monomorphic and megamorphic state.
if (o.getClass() == A.class) {
// inlined variant that matches A
} else if (o.getClass() == B.class) {
// inlined variant that matches B
} else if (o.getClass() == C.class) {
// inlined varint that matches C
} else {
$Deoptimize(); // -> exit, this never returns.
}
[Note: `$GetShape(o)` became `o.getClass()`]V8 can do this. If inlined variants for A, B and C are all the same V8 can also do
// Check that we are either A, B, C
if (o.getClass() != A.class &&
o.getClass() != B.class &&
o.getClass() != C.class) {
$Deoptimize(); // -> exit, this never returns
}
/* inlined variant for A, B, C */
However in both cases you pay penalty for the conditional control flow and in the first case V8 is unable to merge subsequent decision trees in any way to eliminate redundancy between them so `o.x + o.y` might end up compiled into something like (assuming polymorphic code that uses objects of shapes {x, y} and {y, x}): int o_x;
if (o.getClass() == A.class) {
o_x = $LoadByOffset(o, 12);
} else if (o.getClass() == B.class) {
o_x = $LoadByOffset(o, 16);
} else {
$Deoptimize(); // -> exit, this never returns.
}
int o_y;
if (o.getClass() == A.class) {
o_y = $LoadByOffset(o, 16);
} else if (o.getClass() == B.class) {
o_y = $LoadByOffset(o, 12);
} else {
$Deoptimize(); // -> exit, this never returns.
}
o_x + o_y
That's a lot of repetitive branching right here. Arguably the code quality can be substantially improved by merging ifs like this: int o_x, o_y;
if (o.getClass() == A.class) {
o_x = $LoadByOffset(o, 12);
o_y = $LoadByOffset(o, 16);
} else if (o.getClass() == B.class) {
o_x = $LoadByOffset(o, 16);
o_y = $LoadByOffset(o, 12);
} else {
$Deoptimize(); // -> exit, this never returns.
}
o_x + o_y
but at the moment V8 does not do this hence the penalty is higher.