(and if we talk about optimizing compilation things are getting more complicated as there int31, int32 and float64 can all coexist and you can have 31-bit integer stored in a float64 value - if operation was specialized for floating point values)
857 karma · joined September 29, 2010
(and if we talk about optimizing compilation things are getting more complicated as there int31, int32 and float64 can all coexist and you can have 31-bit integer stored in a float64 value - if operation was specialized for floating point values)
That however does not mean you need have to minify your server code.
Please don't base such decisions on some random microbenchmarks that have 0 connection to the code that you are really running in production.
Instead base these decisions on profiling of your actual application.
As I already said below: it is unlikely you get 70% improvement by removing all the comments in your code. For that you need an app that does nothing but a very hot tight loop calling the same small function that does almost nothing too.
JSPerf test case has misleading name: it has nothing to do with inline caching, but it is indeed related to inlining heuristics. V8 uses source size in heuristics for historical reasons and source size includes comments (see the issue above for more info).
Saying that comments slow down your code by 70% is an exaggeration: if you remove all comments from your code it is unlikely that you will get 70% speedup across the board on V8, though potentially you might get some. Test case is really written in a way that highlights the issue, because loop itself is tight, it's doing nothing but a monomorphic function call and target function also is doing mostly nothing - this makes inlining of the callee essential for peak performance here. It is unlikely however that your code contains a lot of loops like this.
Truly efficient FFI has to be part of the VM, you can't implement efficient FFI from outside.
> - I access data (by physical address) allocated externally
Do you really mean "physical address" here?
External data is created precisely to efficiently access data allocated externally, because it allows direct raw access to a given region of memory (sans bounds checking).
Can't you just expose the data you are reading as an external typed data array to the Dart code? That would remove any need for those methods.
> - Dart developers does not like annotation
I like annotations!
> - Dart VM executes custom "native (which in fact are very fast)" methods very slow
This concern is very valid: it is true that transition between Dart code and native methods is too heavyweight and as a developer you have no way to fix it yourself. We had plans to fix it eventually - but they never been very high on the list of things.
Did you file a bug for the slowness of your native extension?
The fear is actually completely opposite: people would misuse these annotations without understanding what they do and then complain that Dart VM "is slow".
As I said: I'd love to convert these lists to annotations. Don't tell anyone but we actually have inlining annotations[1] which we use when writing tests.
> Do you use such a layer as a "Well structured intermediate language" in Dart VM?
I don't understand the question. Dart VM uses intermediate representation when compiling, which is an implementation detail (unlike CLR IL - which well specified input to the VM essentially)
[1] https://code.google.com/p/dart/source/browse/branches/bleedi...)
Annotation like that is not enough for intrinsics. You would still need to provide some way of detecting in the compilation pipeline and intrinsifier that method you are looking at is indeed `Integer_bitAndFromInteger`. In this particular case you could look at the name of the native, but that does not work with those intrinsified methods that have Dart bodies. In those cases it would have to be @intrinsic("name of the intrinsic").
Right now we can write C++ code like this:
switch (recognized_kind()) {
case MethodRecognizer::kSmth:
case MethodRecognizer::kSmthElse:
break;
}
If we start to relying on strings - this code will become less readable (or alternatively you would have to map strings into C++ enumeration which would require a list quite list one of the above).To be honest aesthetically I like annotations, so I would like to use them to replace these lists, but that does not really solve anything or improve much. Even more: there is really no fundamental difference between having an annotation and a list like that, so I don't really understand what bothers you about these lists.
> (dynamism is better, determine everything at runtime or from magic lists)
Annotation would be no less magic (and not available to the user code anyways).
"Using asm.js" is not some magic that just makes all JavaScript faster. For example all those benchmarks from http://dartlang.org/performance are written in a normal JavaScript not in what is known as "asm.js".
Now TurboFan does not even really "use asm.js". It respects "use asm" directive to decide which functions to optimize and enables a few optimizations on such functions. Other than that it's a completely generic JavaScript JIT compiler, which does not even performs a complete asm.js module type verification pass (something that OdinMonkey does).
Furthermore if you try enabling TurboFan on a non-asm.js code then you'll discover that right now TurobFan is lacking infrastructure for adaptive optimizations which is an absolute requirement for compiling normal JavaScript (as opposed to asm.js compliant code which is limted to arithmetic, typed arrays manipulation and essentially statically typed) into efficient machine code. It will run this normal JavaScript code - but it will not necessarily run it fast.
That said I am looking forward for TF with those adaptive optimizations implemented. It's a bold project with a lot of promised value for V8. Though I would not expect anything as impressive as "6x over Crankshaft across the board".
This claim is substantially false. Where are you deriving this "Dart VM has a lot problem..." from? (feel free to respond in Russian). If you actually knew how V8 is implemented you would know that all values in the V8 have something called a hidden class (internally in the sources it's called map) and V8 optimizes based on identity of those hidden classes. This is not that different from the Dart - classes are just not hidden in Dart.
> A lot of (problematic) method in these list. Dart VM code generators and optimizers is unable to optimize usage of them by the Dart VM intelligence.
You are misunderstanding/misrepresenting why these lists exist. Let's talk about them one by one.
1. INTRINSICS Most methods on the intrinsic list don't even have pure-Dart bodies (so there is nothing to optimized with "Dart VM intelligence") and they are on the list for two purposes:
- provide fast hand written assembly implementation, that is more efficient version written in C++; - tell optimizing compiler how to lower these instructions into IR operations.
Without them being on the list there is no way optimizing compiler can optimize anything - because it would only see a call to some native runtime function with unknown semantics. How is it supposed to understand that _TypedList._getFloat64 is actually a pair of a CheckArrayBound and LoadIndexed<Double>() instructions?
Some of these intrinsified methods do have Dart bodies - but we still provide a hand-written assembly implementation for them. I think these days it's limited to methods of the Bigint class. This is the same in any arbitrary width integer library (e.g. look into GMP sources) - you can write it in C++ but people still write these things in assembly because even C++ can't always produce the best code for this.
2. INLINING_WHITELIST These are the methods that that we know are usually beneficial to inline into the caller's code even if other budget constraints are preventing it. Inlining is a very hard problem in compiler construction - it's impossible to always make optimal decisions given compilation time constraints of the JIT. Even AOT compilers need hints for this, and if you take a random JIT you will discover that it most likely has a similar whitelist built in. Essentially this white list is here precisely because we know that we can optimize this methods very well once they are inlined (which is complete opposite of "cannot be optimized" that you invoke).
3. INLINING_BLACKLIST Similar to whitelist it contains functions that we know are not beneficial to inline. Limited to Bigint class actually now, which is a very special case as explained above.
4. POLYMORPHIC_TARGET_LIST this is utility list that improves the structure of the graph that polymorphic inliner builds (it tells inliner that resulting code would not benefit from merging different branches of the polymorphic call even if they share the same target). None of the methods on this list have an actual Dart body.
To summarize all these lists exist because they encode knowledge that is impossible for Dart VM to get because those properties are essentially algorithmically undecidable.
fwiw you can do this on chromium.googlesource.com as well. look for link called `[path history]` when navigating the tree.
btw, if you have a moment I would really appreciate if you tell me how to reproduce +Brandon Donnelson results. It's hard to figure out from those photos which version of GWT should I get from where and what to compile with it. I was not sure if I supposed to check out Joel's code AS IS or I should get it from some other place, etc.
Are you sure? I have never looked at Java implementation of Box2D used in this benchmark, but Dart version is definitely not that polymorphic in its hottest function which corresponds to this Java one:
https://github.com/jbox2d/jbox2d/blob/76fa2602a6abcbc557c9d5...
It looks like mostly a bunch of floating point math to me with not that many calls out.
Another interesting thing I noticed now is that Java version has all vector math inlined manually, while in Dart version it is not the case (at least not entirely if I remember correctly - we want to write high level code and let optimizer do what it can).
Now check out performance graphs of LuaJIT2 compiler+interpreter vs interpreter modes[1].
Anything remotely computationally expensive is 2x faster with compiler and you can go up to 28x for integer number crunching.
I do believe that original point "a good interpreter can get a large portion of the gains you'd get from a compiler" can't be correct simply because it is too broad and ill-defined. What are the gains you expect from the compiler? How can "large portion" be defined? All of these really depend on many things: from the language itself to concrete design decisions in compiler/interpreter.
Also any cross-language comparison should be done very accurately - because we are talking about different language semantics and different benchmark implementations.
Though swiping should have worked too.
[1] I am well aware about different attempts to attack this issue from various angles from AOT to JIT generating JS code but no attempt currently produces truly efficient solution that demonstrates small footprint and consistent high performance on par with native VM across all of the web or even across its relatively modern part.
Microbenchmarks like that require a lot of care to measure something correctly. Check out my talk from LXJS2013 for more details: slides http://mrale.ph/talks/lxjs2013/ and video https://www.youtube.com/watch?v=65-RbBwZQdU
So I always putting this into not-done-yet category for JS VM related work and it is a very interesting problem to tackle.
You are comparing different ratios: munificent was comparing against JS and your "2 times" prediction is against native performance of the compiled VM.
> there are interesting results even there, see pypy.js.
Yes, the most interesting result is that you have to warm up pypy.js with a blow torch, otherwise performance is abysmal.
> I think that approach could work for Dart as well.
Have you seen people complaining about size of dart2js output? Can you imagine how big emscripten output would be?
I have been arguing that dart2js could have been using hand-written JIT compiler on the client side, but the startup performance compared to AOT would really be a big deal. I think a combination of AOT and JIT would be the best, but sadly it is also true that JavaScript lacks right level of abstraction. It is either too low-level (typed arrays, hand rolled allocations etc) or too high-level.
JSC might be warming faster or simply have a better cost / benefit ratio at lower tiers. I doubt FTL kicks at all for this benchmark given how short running it is. V8 also deopts quite a bit on it, so it might be some pathological case... This is definitely something worth exploring and fixing on V8 side.
Care to elaborate on this, what kind of GC overhead are you seeing when creating small object copies?
This is not correct. V8 does sink allocations into deoptimization exits. It does not sink allocations out of the loops at the moment though.
> I think it is unknown how to do this well in method JITs
I don't think it is unknown. The main simplification for tracing JITs comes from the fact that deoptimization and loop-exit can be elegantly treated within the uniform framework, which is a little bit harder for method JIT and you need to find right place to insert materialization instruction after the loop based on post-domination. Nothing hard or unsolvable though.