The Emterpreter: Run asm.js code before it can be parsed
blog.mozilla.org
blog.mozilla.org
I mean, with asm.js and the copious amounts of compile-to-Javascript languages available these days, Javascript is already becoming a de facto bytecode for the web. But it's always been a weird and uncomfortable hack done for the sake of backwards compatibility: a higher-level language hijacked to work as a compile target, simply because it's the only thing supported by browsers.
If projects like this Emterpreter catch on, though, it allows for a smooth path to proper bytecode: for backwards compatibility, you have the Emterpreter read and execute the bytecode, but in other, more modern browsers you have the browser execute the bytecode directly. I think this would be an overall better approach than what we have now.
For large applications compiling to asm.js (such as Unity3D) this experiment could provide significant gains in load time. Considering games usually spend their initial time presenting a menu, background loading of the fast-path makes a ton of sense.
I don't see "load time" as a good enough reason for a revolution of this magnitude. Do you?
1. [citation needed] and a [specific bytecode] needed. Details are everything: an uncompressed .dex file is comparable to a gzipped compile jvm class file; Also, with a specific bytecode in mind, if you compared minified+gzipped gs to gzipped bytecode, I suspect the difference will be small.
2. Regardless, the benefit of binary size is already provided, yesterday, with the emterpreter, without requiring any buy in from any browser vendor.
So, again - what's the benefit, other than startup time (which might be solved in other ways) of a universal browser bytecode?
And even if it never does - it's politically close to impossible to agree on a universal bytecode, so it is extremely unlikely to happen. (Google already tried with PNaCl - if they can't pull it, I doubt anyone else can)
So as things stand right now we have 2 popular "VMs" people really use. One is JavaScript since it's going nowhere. And the other is LLVM that is "free as in beer" for companies all over. JavaScript was secluded to the client. And LLVM was secluded to the backend. Now they are going to be marrying and having lots of children. :-)
A smart compiler could recognize the ordering priority and load in chunks sequentially, with hot code rolled in first, less frequently exercised methods last.
Here's to hybrid vigor!
So free software/open source is analogous to an open society without any arbitrary marriage restrictions, whereas closed source proprietary software tends to cause aristocratic inbreeding?
I suspect that bastards also fit into the analogy somewhere.
It was a stupid email - you can make LLVM into an architecture-agnostic bytecode by disallowing those codes.
Don't believe me?... https://developer.chrome.com/native-client/reference/pnacl-b...
No, I'm not. In fact, I have no idea what email you're referencing.
LLVM as an intermediate language (I guess you could call it bytecode if you really wanted to) for compilers of arbitrary languages (expressly not a virtual machine!) is the only point I intended to make.
http://lists.cs.uiuc.edu/pipermail/llvmdev/2011-October/0437...
It was not a stupid email. The arguments there are still pretty much correct today.
I think that's very unlikely to happen. Instead, what will happen will be exactly "emterpreter"ing: There will be multiple bytescodes, and they will all ship with their own interpreters/compilers.
If the emterpreter, instead of executing the bytecode, would generate JS code for it (and feed it into the JIT compiler if there is one), you'll get the best of all worlds -- bytecode format, top performance, perfect backwards compatibility. I suspect that's kripken's next move.
Given that this is the case, why would someone want to shackle themselves to a specific bytecode format, which is practically impossible to get universally accepted? (This argument is supported by PNaCl; The technical problem is small; the political problem is huge).
Work on JS optimization by all vendors is already extremely impressive and is not going to stop even if everyone agreed on some bytecode. Why not capitalize on it? What does a specific bytecode buy you beyond slightly shorter load times (which the emterpreter already gives a way to greatly reduce), and not having to pull the (essentially universally cached) emterpreter code?
Because these two things, while nice, are not enough support for the revolution that a proper bytecode is.
If code using this method clearly labels the bytecode and interpreter and the part that background loads the "real thing", then implementations can opt to add whatever optimisations they like to speed it up as it stabilises . Doesn't matter if the bytecode changes, as long as it's labelled properly so the optimised versions falls back to just interpreting the JS if it comes across a version (of the interpreter/bytecode as a whole, or just a single opcode) it doesn't understand (or that the implementer hasn't seen a need to optimise).
If the interpreter is guaranteed to retain a certain structure, it could be very easy to just "unroll" the interpreter loop and selectively JIT portions of the bytecode based on hotspots. You can optimise that a lot in a non-bytecode specific way by annotating the interpreter loop with assertions that grants extra guarantees (immutable bytecode; markers to indicate which code is only interpreter scaffolding; decoding hints; if you also tack on "labels" for each instructions, implementations can special case on individual instructions that "settle" while still handling new instructions/changes by inlining the interpreter code.
> What does a specific bytecode buy you beyond slightly shorter load times (which the emterpreter already gives a way to greatly reduce)
The full speed from the start; note the substantially lower speed for the first little part. And the example codebase is small compared to some of the things people want to run.
I think we sort-of agree. I don't necessarily think there's a reason to specify a standard bytecode, exactly because this approach could conceivably be extended to effectively give us a "mostly standard" bytecode with the freedom to continue to change the format without a lengthy committee approach, because there's a demonstrably viable fallback.
We already have a proper bytecode with several high-performance implementations that's supported on virtually every platform, which works as an excellent compilation target for both dynamic and static languages.
It's called JavaScript.
Now you might say JS is bloated! But that's only if it's not gzipped or minified.
Make a "proper bytecode" and you gain absolutely nothing except losing the wide support JS enjoys.
For example, a JS parser specially optimised for asm.js.
It is not. Where would we be if we could only execute Java on the JVM? Clojure and Scala transpiled into Java? Can you imagine how terrible that would be?
I can, and it won't be terrible at all. Compile times might suffer a little (but that's likely to be negligible with a good Java compiler - I think Jikes is now abandoned, but back in 2002 it was pretty much instantaneous even for huge files, unlike Sun's). I don't think you'd notice it on Scala, and likely also not in other JVM languages.
I've worked with Python extensions that compile indirectly through C, and I have no experience with Nim but it seems to do that very well and very quickly.
What exactly do you believe the problem would have been if Scala or Clojure generated Java rather than JVM bytecode?
[edit: fixed a formulation which sounded like code bloat can be blamed on emscripten instead of user-code]
Also, what makes people more cautious about performance concerns is that mobile network and hardware are still catching up to what people have on the desktop. And since about 8 years ago, mobile has been a big opportunity for many companies, which means that the desktop has been taken for a ride by the mobile devices and that is not going to stop.
But there are some huge multimillion line codebases that you can't really reduce in size to 5MB. That's the startup time problem that the Emterpreter aims to help with.
I actually find the bytecode part more interesting then the interpreter part, are you seeing drastically better compression for a bytecode module compared to compressed ASCII asm.js?
However, it does seem likely that a binary format that is designed for compressibility could be smaller. But, it would not be as fast to execute.
SDE effectively compresses a low level syntax tree by building a dictionary that as a sort of sliding window over the tree by adding specialised nodes as well as more general nodes, and the encoding new variations of specialised sub-trees using the new elements in the dictionary and subsequently adding more specialised "instructions".
As a result you both get compression (think a tree variation of Huffman coding) and "hints" to let you re-use partially generated code as templates (e.g. one dictionary entry maps to "assign x to y"; you also add "assign x to _", and later you find a reference to the entry for "assign x to _" + z; now you copy the code for "assign x to <something>" into place and fix it up with the address of z).
The approach has always fascinated me, but it fell in the shadow of the JVM.
[1] https://en.wikipedia.org/wiki/Semantic_dictionary_encoding
Non-blacklisted emterpreter looks slow enough (5fps ish on that graph?) to simply not be useful for some use cases, like a game engine - it's not going to be remotely playable like that. Therefore emterpret => asm.js actually significantly increases the startup time. Playable by 1400ms is worse than playable by 700ms. But I guess this is all preliminary and improvable though!
* Compiling asm.js can use multiple CPU cores, so doing just that is faster than doing it while the emterpreter is running on (at least) one core.
* I believe, but am not sure, that compiling on a background thread is done at lower priority than stuff on the main thread.
* Swapping asm.js code in can only be done in between frames. At 10fps for example, that means around a 100ms delay just for that, and possibly more depending on the state of the browser's event queue.
However, he has JS hanging on for longer than I bet on... once asm.js gets a good DOM binding I expect the explosion of language diversity to take about two years, tops, and for it to rapidly become clear that JS is now just another way of accessing the DOM. I think there's more pressure built up there than people realize, because right now there's no point in thinking about it, but once it's possible, kablooie. Node's value proposition, IMHO, is in some sense correct, but backwards; it's not that we want to write in Javascript on the server, it's that we want "client language = server language"... and once there's no longer a technical handcuff pinning the client side of that equation to Javascript, it will not take that long for it to no longer be Javascript. It is not an impressive language, even within its own 1990s-style dynamic language niche.
(I think this is not because it's "bad", but because it has been developed in this really terrible multiple-vendors-that-actively-don't-want-to-cooperate way for most of its lifetime. It's gotten past that, I think, but during those decades all the other scripting languages were marching right along. None of the other languages could have survived such a process and gotten to where they are today, either.)
There's no particular reason why not. It's all just bits and bytes in the end, and asm.js gives a pretty low-level view of the world. And if you're starting from a baseline of a language that can easily be 5-10x faster than browser-based JS you can afford a bit extra on the GC side.
Javascript isn't magic. It's just a language. It isn't even a particularly special one, once you ignore its browser support, and it certainly isn't one focused on performance (I stopped buying the "languages don't have performance characteristics" line a while ago). It gets to run the same assembly instructions everybody else does. It isn't as fast as a lot of people here suppose, and it isn't that hard to beat out its performance even now.
(Sorry, I can't condone the word "transpile". Usage of it just reveals someone who doesn't understand compilation technology and thinks there's somehow something "special" about compiling to one intermediate language ("javascript") vs. another ("assembler").)
I can't believe how many people seem to believe that Javascript is a C-speed level language, and downmod anyone who observes it's not. Well, it's still not. It's easy to see that it's not. It's not even close. If it were asm.js wouldn't exist. (I mean, if you're having trouble with my claim here, stop and think about that for a moment... if Javascript is so fast, why does asm.js even exist?)
I hope we'll eventually see a proper bytecode spec with bidirectional assembly/disassembly w.r.t. JS (ie: more a transformation than an assembly spec) evolve from this effort, but it's obviously something that needs to happen after asm.js has had its time to bake.
This, of course, assumes that browsers don't get AOT/background compilation to the point where it's no longer necessary to consider a bytecode spec.
Why are we reinventing the wheel?
I sincerely hope I will be wrong!
As for why not the Java VM, my guess is that the browsers are staggeringly enormous piles of C++ code and trying to integrate a Java VM into it would probably be insanely difficult, and anything other than pure 100% integration, too slow to use. It is probably literally easier to continue with the already-integrated JS VM and improve it up to JVM-esque quality than to try to graft the JVM into the browsers that exist today.
Or somebody would have already have tried since nothing would have prevented JS from already running on the JVM, if that were feasible; asm.js is actually independent of this question when it comes down to it.
And you know the JVM ran in the browser almost 20 years ago, on hardware with a fraction of the CPU and memory resources we have today?
Yes. So?
> And you know the JVM ran in the browser almost 20 years ago,
And it did such a good job that everyone has been prefering java applets to html/js/flash for those 20 years.
Oh, actually, they didn't. In fact, with few exceptions, they have been rejected by users for most of those 20 years. The implementation was horrible. I don't know if it's better today - maybe it is for some Windows browsers. But no browser on my Mac or Linux supports it without an external install -- and last time I actually used it (~2 years ago), it still took forever to start applets.
It might have worked if the execution was acceptable; it's not impossible - flash had reasonable execution which lead to widespread adoption that is still nontrivial despite a big part of the web - that of mobile - made it unusable. And yet, 18 years later (I'm up-to-date as of 2 years ago), Java applets are still a mess.
So, the bottom line is: who cares the JVM ran in the browser?
Instead, we're developing apps on a platform intended for delivering documents, and targeting code to a "bytecode" (asm.js) built on a half-baked language originally intended for doing form validation, pop-up ads, and animating dancing bears.
Java and the JVM form a 3/4 baked language environment originally designed to run on washing machines, freezers and TV-top sets - that after 20 years can't even get applets and GUI properly.
See how easy it is?
Yes, I completely dislike JavaScript. I also completely dislike Java. But both are here to stay, regardless of their (lack of) merits compared to some ideals. Java missed the browser bus because it was horrible, and JavaScript drives the bus now because it was there.
The JVM has never run in the same process as the browser, to the best of my knowledge, with the exception of the ill-fated "HotJava" browser: http://en.wikipedia.org/wiki/HotJava which ran Java applets "in" the browser by virtue of being written in Java, instead of C++. At the time this was too sluggish for general use, though.
So allow me to repeat myself: We can't run the JVM in the browser right now. The impedance mismatch between the JVM and the C++ world are just too great to have sufficient performance right now. All the current browsers are enormous piles of C++ code. There is no way to "just" integrate a JVM directly into them, and no bridge between the two is going to have enough performance for the demands we're putting on browsers. It doesn't matter how awesome Java may or may not be when the browsers aren't in Java. It's not an option today, short of replicating Sun's feat and rewriting your own browser in Java. But the only thing stopping you from doing that is the sheer size of the task, rather than any technical problem. (But make no mistake, it is an enormous undertaking now to write an engine capable of replacing any of the existing ones for even a single well-chosen use case.)
* Patent and copyright issues - the lawsuit with Google is still going on, last I heard.
* There used to be serious technical issues with startup speed. People saw websites with Java applets and saw how slow they were to load. If that hasn't been fixed, it's a serious problem, as websites do need to load fast, unlike typical Java applications.
edit:
* A big use case is compiled C++ code. I am aware of lots of languages compiling to the JVM, but I actually don't think I heard of C and C++. Is there such an option? If such an option doesn't exist, or exists but runs more slowly than asm.js currently does (which is pretty close to native already), it would be a problem.
Sure there's some parallels, but don't pretend industry ditched Java applets due to ignorance. Javascript in the browser, for 99% of web applications, provides a much better user experience. The developer experience we could argue here, but I'd say nowadays developing client side JS (or TypeScript, CoffeeScript, Haxe etc.) isn't that bad. The tooling is great and there's if anything, too many frameworks to choose from (competition is good!).
The same arguments apply for Flash, really. Flash integrated slightly better with browsers at the expense of being a resource hog.
Perhaps the browser/OS vendors and Oracle could have worked together to make this happen, but they didn't, and the world moved on.
Anyway, plugins are dead code walking, not only because of security worries, but because "mobile" in full trumps desktop, and the de-facto "no plugins on mobile" OS-set standard.
Plugins alas are still "facultative not obligate" on the desktop, so browsers support them. Even on desktop, their days are numbered (see Chrome killing NPAPI plugin support; see also Shumway).
The JVM isn't going to be integrated directly (i.e., not as a plugin) into browsers, either.
For all its virtues on the server side, and I'm thinking of multiple JVMs here going back to the CLDC VM that Macromedia tried to license from Sun for Flash (rejected: they did Tamarin instead), "the" JVM is nowhere near a clean fit on the client side.
This leaves the already-obligate-not-facultative JS VMs, which are evolving fairly rapidly due to browser competition, serving both hand-coded and compile-to-JS workloads.
All as predicted! I think Dave Herman first spoke about this at Web Rebels in 2012, and a bunch of us have laid bets well before then, including on HN. Yeah, I've been hard on PNaCl as a would-be better plugin for safe native code, not due to the tech itself as because of the opportunity cost.
Now that it's 2015, everyone seems to be on convergent paths toward some kind of LLVM-compiled, cross-browser, safe-native intermediate code that's evolved from JS (starting from asm.js, but not restricted to the current subset, e.g., shared memory threads could be added only to the intermediate code and its runtime).
Such a "WebAsm" or js.bin format would then co-evolve with JS until source code download goes away -- if JS source download ever does (I'm skeptical, but given enough time, it could happen).
Lots of risk still, but this remains by far the shortest-distance evolutionary path from where we are.
/be
Two words: "Java applets".
Seems that caching commonly shared dependencies could be a good way to cut down on size and parse/compile time.
After a while, if a chunk of code is identified as "hot", a second-stage compilation kicks in. The code is recompiled to optimized machine code that takes advantage of the types that it previously saw the code use (with fallbacks in case those assumptions later fail).
it would effectively make a js -> asm.js compiler
Using previous runs to inform the JIT of expected types is entirely reasonable though, and think various JS implementations already do this.
none of that should matter. if the resulting JIT'd code is faster with the traps than the plain js without them, so be it.
I don't know how desirable it would be - you'd be shifting the burden from the compiler to the network, basically.
In particular, you'd still have to compile it, which means that all you're gaining is the time before JIT decides that the code's super hot.
A nicer approach might be like Oracle hints - you have a small comment you place above a function that tells the JIT how you want it compiled. You test your software using an instrumented browser, and that adds the comments between you finishing and it running.
It kinda ups the complexity by loads, though - and I bet the benefits would be relatively small. Stuff like minifying and compressed assets has really tangible benefits, but here you have something that can easily be done wrong, and really only improves the execution of a very small overall percentage of functions.