https://en.wikipedia.org/wiki/UNCOL
Outside the browser it adds very little over Java, MSCLI, LLVM, P-Code, M-Code, Python Bytecodes, TIMI, DEX and a myriad of other formats.
Cortex M4 is already targeted by MicroEJ and MicroPython.
https://en.wikipedia.org/wiki/UNCOL
Outside the browser it adds very little over Java, MSCLI, LLVM, P-Code, M-Code, Python Bytecodes, TIMI, DEX and a myriad of other formats.
Cortex M4 is already targeted by MicroEJ and MicroPython.
No doubt there have been plenty of those.
> Outside the browser it adds very little over Java, MSCLI, LLVM, P-Code, M-Code, Python Bytecodes, TIMI, DEX and a myriad of other formats.
It does add one very important thing: a large ecosystem. One where security and performance matters: the web.
Regarding libraries, graphical debuggers and reverse code generation.
Security and performance were part of Java and .NET design, hence why they have validation as part of their execution workflow.
As for security on the Web, cross site scripting and WebGL exploits prove that is still room for improvement.
I wonder if the real difference between WebAssembly and the other bytecodes you mentioned is just the number languages that will end up targeting it.
Though, one advantage over LLVM is that it's simpler.
Java and CLR bytecode verification is crap because they weren't designed for efficient verification from the start, WASM was: they're expensive, slow, and imprecise. WASM verification gives you much stronger properties, and it's significantly cheaper. It's a much better universal bytecode than any other offerings currently available.
Also WASM remains to be battle tested regarding exploits in the wild.
The difficulty of full safety verification on the JVM is well studied [2]. The security-focused Joe-E language actually decompiles JVM bytecodes back to Java to avoid the many full abstraction failures that have been documented.
The CLR is a little better on the safety record, but the verification costs can be even higher because the CLR supports various pointer types. Proper verification requires control-flow analysis, but instead they partition the bytecode into safe/unsafe variants and only the safe variant is "verifiable". IIRC, WASM doesn't have this limitation.
[1] https://www.cl.cam.ac.uk/~caw77/papers/mechanising-and-verif...
[2] http://www-master.ufr-info-p6.jussieu.fr/2005/IMG/pdf/leroy-...
As for the links I will check them later. Thanks providing them.
WASM still needs to prove itself in similar scenarios.
"Fine-tuning security" is typically symptomatic of bad security design. Security is not a separable concern.
I will read the paper later on.
Modern JVM bytecode is cheap to verify and gives you very strong properties like memory and type safety, i.e. the app will not have buffer overflows or type confusion attacks in it.
WASM gives you virtually no guarantees about the software running in it, beyond that it (maybe) can't escape its sandbox.
It's a much better universal bytecode than any other offerings currently available.
I think that's a very strong statement for something so debatable.
WASM appears, to me eyes, to have numerous serious flaws that make me wonder why people are so excited about it. JVM bytecode appears to beat it in every aspect.
1. No GC support, no real workable plan to get there because there's also no type system worth a damn. They've been looking at it for years and this article says all they've got is a tiny stepping stone - no actual GC, just a way to mark pointers to things that came from JS in the type system. This wouldn't matter if the future of the web was software written in C, but if it is, god help us all.
2. No threading. JVM has been thread safe from the start, and has had a huge amount of work put into its memory model, so you can reason about the nature of bytecode when run on different CPUs and in the presence of multi-threading. Moreover making a runtime like V8 fast is much harder if you try to make it thread safe. JVMs have a long history of being fast in the presence of large scale threading but no WASM supporting VMs do.
3. No support for exceptions, apparently at least one failed attempt to add it. Even LLVM has support for exceptions. Again, no big deal if the future of the web is C99 but what a bad joke if it is.
4. No FFI of any use.
5. No dynamic code loading.
6. Tooling is poor or non-existent.
7. Ignoring these differences, WASM bytecode looks a lot like JVM bytecode e.g. is a stack based language that requires compilation client side via JITC to approach good performance.
As far as I can tell it's a significant regression from what Java could do even 20 years ago. And don't start talking about security. WASM integrations already opened up critical security bugs in browsers:
https://bugs.chromium.org/p/chromium/issues/detail?id=729991
https://bugs.chromium.org/p/chromium/issues/detail?id=794091
https://news.ycombinator.com/item?id=15712270
If you think WASM is magically immune to sandbox bugs, you're wrong.
The reason to be excited about it is to have a universal sandbox to portably run untrusted code regardless of the source. That's unprecedented flexibility for something so widely deployed.
Neither the JVM or the CLR provide such a sandbox for code that uses pointers. LLVM IR is not portable, is not sandboxed and is always in flux.
As for your list of "flaws", they aren't flaws at all. A flaw is a feature that cannot be supported, even in principle. Everything you list are possible as extensions to the core type system.
WASM was designed with efficient verification in mind, and with mechanized formal proofs as I point out at [1]. Java's verification was always a hack, and the mechanized proofs of Java's verification procedure never encompassed all JVM bytecodes, and full verification was always too costly to perform at runtime. Java has numerous security problems, not just with verification, but also vulnerabilities due to full abstraction failures.
As I mentioned in [1], the security-focused Joe-E language explicitly chose to decompile JVM bytecode to plain Java to avoid all of the numerous documented full abstraction failures of the JVM [2].
The similarities between the stack-oriented bytecodes are entirely superficial.
Actually, the CLR supports pointers and the JVM can run arbitrary LLVM bitcode (so C, C++, Rust etc) in a memory safe way with bounds checking and garbage collection these days. Check out Sulong and Safe Sulong:
http://ssw.jku.at/General/Staff/ManuelRigger/ASPLOS18.pdf
JVMs can also manually allocate memory using the Unsafe class, I suppose if you wanted WASM like protection semantics that API could be constrained to a particular memory region and become "SortaUnsafe" for example. You could then compile C to such a dialect. The LLVM bitcode on Graal approach is likely to work better though,.
However, realistically most new software is not being written in C or even C++ for that matter. Most developers use managed languages. And the JVM can do a much better job of that than WASM can.
A flaw is a feature that cannot be supported, even in principle.
This is a fascinating definition of flaw that I haven't previously encountered.
In what sense does the JVM not do "full verification at runtime"? Also if you have some more material on which bits of the JVM bytecode set aren't formally verified I'd like to read that. I can imagine, that the latest features may not have been done, but the older JVM bytecode sets were formally verified at least.
I said it doesn't provide a sandbox for pointers. CLR pointers can corrupt your whole VM instance. WASM's lightweight bytecode with no runtime can support in-process heap isolation.
> JVM can run arbitrary LLVM bitcode (so C, C++, Rust etc) in a memory safe way with bounds checking and garbage collection these days. Check out Sulong and Safe Sulong:
Nice find. Still, look to be interpreted though.
> In what sense does the JVM not do "full verification at runtime"?
The last report of a formalized semantics for verification that I saw was over 10 years old:
https://repositories.lib.utexas.edu/handle/2152/2763
The JVM bytecode has significant verification challenges as discussed in this paper and by Leroy in the other papers I linked. The Joe-E papers I linked also discuss the problems of targeting the JVM at length.
As I linked in my other comment, WASM was built with mechanized proofs nearly from the outset. They learned from the mistakes made in the CLR and JVM.
It gets JIT compiled to native code that runs (for some code shapes yadda yadda usual story) about 10% slower than gcc, if I recall correctly. Without the safety aspect it can run as fast as GCC for some benchmarks.
As I linked in my other comment, WASM was built with mechanized proofs nearly from the outset. They learned from the mistakes made in the CLR and JVM.
But so did the CLR and JVM designers.
The paper you link to says the hard part of JVM bytecode verification is the jsr "subroutine" control flow instruction. This instruction has been phased out, bytecode version 51+ doesn't allow it anymore. Whilst a JVM may well support verification of older bytecode formats, and that verification is still safe and sound, you can implement the latest version of the spec and avoid that entire design error.
Given that, and given that the JVM type system has been proven sound despite being significantly more powerful than WASM's, I'm not really certain how this is meant to prove that WASM is better.
No they weren't. Formalization and verification of the IR came much later in both cases.
> Given that, and given that the JVM type system has been proven sound despite being significantly more powerful than WASM's, I'm not really certain how this is meant to prove that WASM is better.
"Significantly more powerful" is way overstating the case. In fact, I'd hazard that it's flat out wrong. You can easily express some patterns in JVM bytecode, but other patterns not at all.
Secondly, WASM's primitives types are much more flexible than those available on the JVM. For instance, unsigned types.
Finally, the JVM carries a lot of baggage, a) in terms of backwards compatibility as you briefly mention, b) in terms of unnecessary control flow instructions which unnecessarily complicate verification (like exceptions), c) unconfigurable runtime that's poorly suited to some programs, d) an overly complicated security model that's an impediment more than a help, etc. I could probably come up with a dozen more reasons, but that's just off the top of my head.
AFAIK there is no Java-Bytecode involved for this and to optimize the AST-interpreter to something faster you need the Graal-Compiler, most JVMs can't do this. There is no C++/LLVM/Rust to Java-Bytecode compiler that is widely used, while this exists for WASM. Even a huge Codebase like AutoCAD was compiled for the Web. Java-Bytecode wasn't designed for this, you could theoretically do the same with Java-Bytecode but it would probably be much slower. Simply because Java-Bytecode wasn't designed for this purpose.
Java-Bytecode is notoriously hard to verify, WASM is more strict and therefore verification is easier. (With verification I mean the process in the Browser/JVM that checks if the bytecode is actually valid). WASM only allows structured control flow, no arbitrary gotos like Java. This makes the compilers job much easier, many JVMs actually just bail out on irreducible loops and therefore such code might not get optimized. OTOH WASM-Bytecode wasn't designed for interpretation.
I don't know why there is this heated discussion. WASM had the chance to learn from the mistakes in Java's Bytecode and both bytecodes were designed for different purposes after all. Java Bytecode was designed for Java, WASM for being a language-agnostic compilation target. Supporting GC in Java-Bytecode for example is much easier than having a similar thing in WASM.
I don't think JVMs where without any security bugs, since WASM is mainly used in browsers right now, security is a MAJOR concern. Java-Bytecode was designed for a different purpose. Try to compile a language with multi-inheritance to Java-Bytecode, it's a pain. WASM is designed in a more language-agnostic way, that's why GC and Exception Handling is much harder to specify than in the JVM.
BTW: GC'ed languages can already be compiled to WASM: they more or less "just" need to compile the GC to WASM too. The GC proposal would only allow to reuse the embedder's (most likely the browser) GC.
But. Then you admit that compiling Java to WASM is hard because there's no GC, you'd have to ship your entire GC with the app and it wouldn't be fast because GCs like to integrate with the JIT compiler and WASM is the JIT compiler here. Also WASM JITs don't really understand managed code patterns well.
So how is WASM more language agnostic? Seems like (if we ignore Graal) WASM has trouble with languages JVM bytecode is good at, and vice-versa. Quite comparable, no?
BTW even saying C++ is easy to compile to WASM is tricky because WASM apparently has no exception support, and exceptions are a part of the C++ language. Likewise for vendor extensions like vector intrinsics, inline assembly etc. At best you can handle a subset of the language.
So in the end I don't buy that it's really more generic. As munificent says, it seems more like people aren't really sinking their teeth into the tradeoffs involved.
Even if we would agree that the JVM is as good as WASM as a language-agnostic bytecode, WASM still makes sense since it doesn't come with all the baggage of the JVM like class files, many bytecodes that exactly match the Java semantics but can't be used in other languages. Browser-vendors would still have to add new bytecodes for common operations for both size and speed reasons. So it made sense to design a new bytecode format. WASM even allows streaming compilation: The browser can start compiling bytecode before it downloaded the whole file.
Yes, there are a few features missing from WASM. But just look how many applications have already been compiled for the web. The missing features are not that relevant for many large existing and performance-sensitive native applications written in C/C++. We don't need to rewrite this applications in JS. I mean JS wouldn't even be fast enough for that anyways. That's what WASM was designed for and even according to you WASM is better suited for this than Java-Bytecode.
Inline assembly and vector extensions would also be problematic in Java Bytecode.
There is a reason why every "C on the JVM" project, of which there have been many, used MIPS GCC and then interpreted the MIPS machine code at runtime. (Yes, really.)
> No threading
> No exceptions
Along with not forcing any object model, this looks like very reasonable design choices to me. (I'd add "no global mutable state", but it's being added.)
Reason 1: this is very similar to what a (RISC) CPU would offer. If anyone need to implement these features, they can be implemented on top of WASM, as they can be implemented on top of a CPU IS. A small specification not loaded with such complex issues as GC or object model is much easier to keep correct and fast. Also, any advances in e.g. GC can be pushed to older WASM implementations, because they don't replace the VM, they are libraries.
Reason 2: Using message-passing between processes ("workers") instead of threading, and Option/Maybe instead of exceptions, is a known way to avoid a large number of problems. The current WASM design makes this approach simple and reasonably natural. There's a reason why Erlang is designed as it is; WASM can take a page from its book.
Does WASM have in place any mechanisms to prevent Oracle or other malicious destructive companies from co-opting it, copyrighting its APIs, and running it into the ground for profit?
That's not very important though, because those had a 20+ years headstart.
WebAssembly is going to have much much bigger momentum than them (except Java and .NET), to the point that it will eclipse the ecosystem all of them together in 5-10 years.
Using LLVM is exactly what Google tried with PNaCl. It had a lot of drawbacks; WebAssembly is a much improved iteration of this idea.
Microsoft's original plan for CLR was to replace the complete Windows stack with .NET, they just failed to do, because not only they had to deal with technical issues, there were the internal wars from WinDev making sure that would never happen, hence the Longhorn failure, followed by the same ideas reborned as COM on Vista, rebooted as WinRT/UAP/UWP a couple of years later.
Microsoft even had something like LLVM, done on top of the CLR, called Phoenix. It just never came out of MSR into production.
https://en.wikipedia.org/wiki/Phoenix_(compiler_framework)
https://www.infoq.com/news/2008/05/Phoenix-Compiler-Framewor...
Also, GraalVM will happily consume LLVM bitcode.
Standard C++ compiles just fine.
There is no different than using language extensions on GCC, clang or any other C++ compilers.
You are not forced to use GC beyond the interop to other .NET languages.
I have integrated quite a few C++ libraries this way, given that my C++ experience makes it easier than having to deal with P/Invoke attributes and possible linking issues.
WASM exists mostly for political reasons as far as I can tell: it's something people who work on browsers can "own", vs outsource to other older teams with more historical baggage and different backing corporations. I mean, for WASM to be technically compelling you have to buy the idea that the best way to move the web forward is to enable applets-written-in-C without any useful form of DOM interop or GUI. That seems like a rather implausible claim.
As for embedding a JVM, sure, why not? Good JVMs are open source these days and HotSpot starts in 50 msec or less. V8 is very comparable, tech wise in terms of its approach. Lots of code out there targets JVM bytecode.
There are very good technical reasons for WASMs existence. JVM-Bytecode for example is very Java-focused, WASM doesn't need class-files or many bytecodes like invokevirtual. For supporting C++/Rust it doesn't even need GC.
Bringing large applications like PSDFKit, AutoCAD or games into the browser is certainly getting the web forward. Sure, WASM isn't intended for manipulating the DOM. This is still JS's job (at least for the foreseeable future).