WebAssembly Stack Machine
docs.google.com
docs.google.com
https://github.com/WebAssembly/design/issues/755
AST was a good idea and parsing as validation was excellent. Having a general stack machine makes validation much more difficult.
This is a bad decision (stacks vs AST or register) made for the wrong reason (code size). That's a 1990s design assumption when memory was expensive and bandwidth was dear. Now memory is cheap; bandwidth is phenomenal; and latency is expensive.
The stack-machine rules were indeed added late, and it's reasonable to ask whether they might have been improved if there had been more time to iterate. It's also reasonable to ask whether a simple register-machine design with a compression layer on top would have been a better overall design.
However, code size is important for wasm. Smaller code size means less to download between a user clicking a link and viewing content. Networks have gotten faster on average, but bandwidth still matters in many contexts.
PS.
I am talking about problem of trust, which this binary format creates. Here, in Linux, we are solving problem of trust using distributions, maintainers, signed packages, signed repositories, releases. It's why I will trust binary packages from my distribution but will not trust webasm binaries.
A binary encoding does not contribute significantly to obfuscation when it can be trivially undone. WebAssembly is an open standard, and browsers supporting wasm have builtin support for converting it to text and displaying it.
Compiled code can be much harder to read than human-written code, though this is mainly because of lowering and optimization, rather than the final encoding.
WebAsm creates problem. Same problem as Java, Flash, Unity, PNaCl and dozens of other platforms to execute binary blobs from untrusted sources. And the only solution is to add trust, e.g. by publishing heavy-weight libraries for review and patching by third-party maintainers, e.g. such libraries as SDL, game engines, GUI, databases, etc. Otherwise, we will have same situation as with other libraries, e.g. JQuery, when sites are using old version of common library with known security problems for ages, despite that fixed version is freely available.
Memory is still expensive:
* Spreading things out in memory more causes more cache misses, which lowers performance.
* Using more memory increases page faults, which lowers performance.
* I believe using more memory drains batteries faster on mobile devices, but I'm not sure exactly why. Maybe the effort spent shuttling stuff from virtual memory into real memory on page faults?
* On embedded devices, using more memory means you need to have more RAM, which increases the per-device manufacturing cost.
* Using more memory to represent code increases memory pressure, which leads to more GC cycles.
There's also the issue of additional latency from tertiary caches, like the disk, when you are under memory pressure.
Why? This document suggests that it's not that different.
"The main new changes to verification are:
- All branches to the end of a block must have the same arity and same types, including the implicit fall-through to end.
- The true block of if-end constructs must leave the stack at the same height as when the if was entered.
- The true and false blocks of if-else-end constructs must leave the stack at the same height with the same types."
That's easy to 'document' and harder to do. It's what the JVM did, again, in the 90s. Now the verifier chapter in the JVM spec is now about 160 pages long.
(Those problems and their solutions haven't been documented yet AFAIK.)
Because you can also easily transform an AST back into code to reverse engineer?
Reverse engineering code for a stack machine is quite a bit more annoying.
push x
push y
add
Evaluate this symbolically and you get (add x y) naturally.Writing a simple interpreter, I prefer register or stack machine over AST walker. It'll be faster, for one. And there's a chance of interpreting it directly with a loop and switch, without a deserialization step.
Many simple analyses of an AST have an equivalent stack machine form. If the analysis can be done as a post-order traversal (like type-checking, constant folding etc.) you're good to go.
I also don't agree re expensive validation; I think you're wrong. Interface with APIs much more problematic than pure computation, which is not hard to validate. Interface safety is largely the same problem as with plain JS, as I see it. I don't see stack machine vs register machine making almost any difference at all.
It would be interesting to be able to take code compiled with Emscripten for example and run it as part of JVM applications, similar to what NestedVM can do.
In any case I don't have time to implement all of that but it seems like it could be any interesting way to run non-java libraries on the JVM
The question would then be, is the mapping straightforward or you do need to reimplement everything from scratch?
EDIT: I didn't think I would be excited about WebAssembly but I'm now really curious to see what happens with it.
what you say about app development being unified is already true: see atom, steam and all others who embed webkit or similar to create a desktop UI.
Its funny how most of the WASM tools still use s-experssions as the text format. How does that even work, now that the AST is history?
You can think of the old s-expression format as a language that compiles into wasm, a language that is an AST and that happens to have the same types and operations etc. as wasm.
The AST format could open up new possibilities for software, some of which are observable in Lispy languages like Scheme (I won't list them here). Instead, we're looking at locking the software world back into this 1960's model for another 50 years out of a misguided concern for optimization over power.
It's like foregoing the arch because it's more work to craft, and instead coming up with a REALLY efficient way to fit square blocks together. Congratulations, we can build better pyramids, but will never grasp the concept of a cathedral.
To really grasp my point, I BEG you all to watch the following two videos in full and think hard about what Alan Kay & Douglass Crockford have to say about new ideas, building complex structures, and leaving something better for the next generation:
https://www.youtube.com/watch?v=PSGEjv3Tqo0
As Alan Kay states, what is simpler: something that's easier to process, but for which the software written on top of it is massive; or one that takes a bit more overhead, but allows for powerful new ways to model software and reduce complexity?
I believe that an AST model is a major start in inventing "the arch" that's been missing in software, and with something that will proliferate the whole web ... how short-sighted it would be to give that up in favor of "optimizing" the old thing.
Imagine if instead of JavaScript, the language of the web had been Java? Lambdas would not be mainstream; new ways of doing OOP would not be thought of; and all the amazing libraries that have been written because of the ad-hoc object modeling that JavaScript offers. Probably one of the messiest and inefficient languages ever written, yet one of the most powerful ever given. C'mon, let's do it a step more by making it binary, homoiconic, and self-modifying.
Thanks.
Then PNaCl came along with a platform-independent bitcode format based on LLVM IR, which was translated to host's native code in the browser.
Then WASM (also platform-independent) came along striving to be a multi-vendor solution. Unlike the other two WASM directly targets the JavaScript engine. It started out as a serialization format for JavaScript AST.
This might help you:
WebAssembly is very different, and in Chrome it is implemented on top of TurboFan, V8's code generator, rather than Native Client.