The best WebAssembly runtime may be no runtime
00f.net
00f.net
https://hacks.mozilla.org/2021/12/webassembly-and-back-again...
> gVisor requires a platform to implement interception of syscalls, basic context switching, and memory mapping functionality. Internally, gVisor uses an abstraction sensibly called Platform.
Chrome sandbox: https://chromium.googlesource.com/chromium/src/+/refs/heads/...
Firefox sandbox: https://wiki.mozilla.org/Security/Sandbox
Chromium sandbox types summary: https://github.com/chromium/chromium/blob/main/docs/linux/sa...
Minijail: https://github.com/google/minijail :
> Minijail is a sandboxing and containment tool used in ChromeOS and Android. It provides an executable that can be used to launch and sandbox other programs, and a library that can be used by code to sandbox itself.
Chrome vulnerability reward amounts: https://bughunters.google.com/about/rules/5745167867576320/c...
Systemd has SystemCallFilter= to limit processes to certain syscall: https://news.ycombinator.com/item?id=36693366
Nerdctl: https://github.com/containerd/nerdctl
Nerdctl, podman, and podman-remote do rootless containers.
GCC and clang could implement their own bounds checking rules, but C -> WASM -> C is actually <C | anything> -> WASM -> <C | anything>
Source: every language vm under the sun (CLR, JVM, Neko, etc.)
I'm sure it's not a cure all, adds overhead and not applicable in all cases, but every small addition to security wouldn't be a bad thing. I don't see any reason why every command line utility in Unix based OSes, for example, couldn't be sandboxed. Think like wget or curl for example.
A seccomp classic sandboxed process will be at least as secure as any WASM runtime, no matter the engine. Even though the former is running untrusted native code, the attack surface is much more narrow and transparent.
It sounds like you're implying that C coders will opt out of the sandboxing provided by the OS, but that's not possible without coding kernel level code. For userland processes, the sandboxing isn't optional, and your process will be sent a SEGV signal if it tries to access memory it's not allowed to access.
Having worked with asm.js, NaCl, PNaCl (plus a couple other long forgotten competitors like Adobe's Alchemy/Flascc) and finally WASM: WASM is the natural result of an evolutionary process, and considering what all could have gone wrong if business decisions would have overruled technology decisions, what we got out of the whole endeavour is really pretty damn good. It's really not surprising that Google abandondend PNaCl and went full in on WASM.
WASM is superior in pretty much all ways to NaCl other than being a bit more time to go through the whole specification and implementation process. NaCl was a great prototype, and WASM is a much more polished result because of it.
What kept NaCl from being independently implemented? And what did WASM learn from NaCl?
Edit: ah, someone already mentioned NaCl being CPU-specific.
I do understand why people wanted something different and rejected (p)NaCl. But the reason is not technical, it's political (in the broad sense). Any technical issues could have been solved by fixing and extending (p)NaCl, but everybody involved understood it was not going to work
pNaCl's biggest mistake was using LLVM bitcode as a wire format (see, for example, [1]). Another problem was trying to use glibc, instead of newlib. That resulted in a lot of low-level incompatibilities, as glibc itself is too target-specific. And implementing shared objects between portable and native world was really messy.
asm.js appeared as a simpler alternative to NaCl, and then it was quickly replaced by wasm, developed mutually by Mozilla, Google and others.
1. https://groups.google.com/g/native-client-discuss/c/h6GJ8nQd...
There were other good reasons why that wasn't ultimately the road taken, but they largely were know technical reasons.
Either make something so simple that it obviously has no flaws or make something so complex that no flaws are obvious.
1) Web UIs with DOM access from wasm/threads and all that goodness. I want to write my entire web application as a Rust wasm application without thunking through JS.
2) A native application with dynamic plugins as WASI libraries. Writing CLI tools and desktop applications in Rust with a practical method of loading plugins dynamically is .
Everytime you use yew.rs or some other "pure rust" or whatever framework, it's making tons of JS calls under the hood.
This is OK for some purposes, but if you're trying for performance on a workload that isn't CPU-bound, it sucks. And it's another abstraction layer that's bound to leak, I imagine.
This is a significant issue with WASM that a lot of us are waiting with bated breath to be fixed. We want to really be able to target the web browser with NO javascript involved.
And that's a good thing. Both APIs (DOM and to a lesser extent WebGPU) are built on top of the Javascript object model. Trying to map the JS object model to WASM inside the WASM spec would be a really bad idea.
> This is a significant issue with WASM
It's a non-issue really, brought up again and again by people who didn't really look into WASM in detail yet and what it is good for (hint: not to replace the HTML+CSS+JS combo).
The DOM is so slow to begin with that going through a (potentially automatically generated) Javascript shim from WASM wouldn't move the performance needle.
For APIs like WebGPU it would make more sense to have a WASM-friendly API, but that's really a design wart of WebGPU's Javascript API (WebGL actually did a slightly better job there). And besides, the overhead of calling from WASM into web APIs is mostly negligible compared to the overhead that happens inside the API (speaking specifically of WebGL and WebGPU here from my own experience).
For one thing, when a browser does the initial load of the page html, it scans for <script> tags, downloads all of them in parallel and does some preemptive compilation while preserving the declaration order for actual evaluation.
WASM modules do not receive that same optimisation because they are loaded during JavaScript evaluation, meaning the browser cannot know they exist until the JavaScript is evaluated forgoing their ability to be optimised ahead of time.
Perhaps some level of optimisation is available through the use of service worker, but it's unlikely that the initial load will be as fast in this context as there is additional overhead associated with establishing a SW and a SW only interacts with the network layer so AOT compilation optimisations will be unavailable.
2) Tooling
One of the areas of promise for WASM driven web applications was being able to exclusively use the tooling of the language you're consuming. The reliance on JavaScript for WASM modules inherently relies on the wild world of JavaScript tooling.
Personally, I have many esoteric use cases that require access to DOM APIs but no access to GUI (third party scripts sandboxed within an iframe for consumption). These use cases _must_ be optimised for initial page load. I have experimented with using Rust for my use case and, while it is faster during runtime, the initial load is slower.
3) Extras I'd like to see
I would love it if I could embed html (as a sort of manifest) within the wasm module such that the browser could use wasm as an entry point - removing the need to load html first, then wasm. As good as http2 multiplexing is, for whatever reason, loading html +1 file is slower than loading html with the application embedded within the file.
I'd also like to see us eliminate the need for the `Cross-Origin-Opener-Policy: same-site` and `Cross-Origin-Embedder-Policy: require-corp` headers as they make multi threading in the browser largely impractical.
...not sure what you mean with that, but WASM isn't actually AOT compiled in browsers anymore, instead there are several tiers of JIT-ting on a per function level. E.g. when a WASM function is only called once it is compiled quickly but without optimization, and then when called more frequently, higher compilation tiers will kick in which compile the WASM function in the background with more optimizations.
This is pretty much the same strategy as used with Javascript code.
You can also split your WASM code into several dynamically loaded modules (and if neeeded, loaded and instantiated in parallel), but (AFAIK) unlike with Javascript bundlers this can't be done automatically. You'll have to design your code that's compiled to WASM from the ground up for being split into several modules (similar to how DLLs are used in native code).
> The reliance on JavaScript for WASM modules inherently relies on the wild world of JavaScript tooling.
I write all my WASM code with C/C++ tooling only, no npm or similar involved. There's only a minimal .html file which defines how a WebGL/WebGPU canvas integrates with the browser environment.
> I'd also like to see us eliminate the need for the `Cross-Origin-Opener-Policy: same-site` and `Cross-Origin-Embedder-Policy: require-corp` headers as they make multi threading in the browser largely impractical.
This I fully agree with. There is a "legal" client-side workaround using service workers btw: https://dev.to/stefnotch/enabling-coop-coep-without-touching....
Why doesn't HTML get an overhaul, with a massive set of new and actually useful and styleable components (e.g., data table with sorting/paging/filtering capabilities built right into the browser), ability to scope CSS, and things like that? That's how the shittiness of web dev can actually be solved so that reliance of client side code can be drastically reduced
For instance, personally I really appreciate that I can bundle an assembler written in C89 in the mid-90's and last updated in 2008, and a home computer emulator written in C and C++ into a VSCode extension without having to worry about platform compatibility or security issues (https://marketplace.visualstudio.com/items?itemName=floooh.v...). Stuff like this wasn't really possible until asm.js and then WebAssembly.
It doesn't, though, which is a big problem. WASM starts the heap at 0x0 which breaks the null handling of 99% of traditional AOT languages _and also_ breaks the null handling of optimized higher level runtimes that rely on page faults to identify that a null deref happened before backtracking and converting that into an exception.
The MVP priority for WASM was to protect the host and make it simple to be the host. The guest features, security, and general sanity is very poor as a result.
Programming languages and their runtimes should not depend on such hardware features IMHO (at most use them optionally for performance optimization).
Also, my guess is that WASM engines outside the browser would absolutely be able to trap on "zero page" accesses.
AFAIK the limitation of address zero being a regular memory location when running WASM in browsers is just a side effect of the WASM heap being a regular Javascript ArrayBuffer object, which cannot start at an index greater than zero (and adding an offset for every heap access would be prohibitively expensive). A WASM runtime which doesn't require to interoperate with Javascript could map the WASM heap to regular virtual memory with the first page being read/write protected.
For higher level runtimes, can you name some common ones that use page faults that way? I've never heard of that technique.
The language doesn't make any promise yet on every major platform for the last 40 years a null deref resulted in a segfault because the lower address space is marked invalid.
This is absolutely, 100% baked into the practical expectations of every traditional AOT language at this point.
> For higher level runtimes, can you name some common ones that use page faults that way?
Android's ART does. I'd be shocked if OpenJDK doesn't for the same reason. Same with .NET or any other mature JIT runtime. It's obviously ridiculously expensive to insert null checks before every object access, and since the CPU & some page protections can do it for free why would you pay that?
This is actually the other way around. Both .NET and JVM do it the same way - they inject the smallest instruction that dereferences a pointer (e.g. cmp byte ptr [rax], al or ldr xzr, [x2]) and when it throws a hardware exception, it is then caught by the registered runtime handler, which then subsequently resumes the execution of managed code from the hardware exception location and raises corresponding Null Reference/Pointer exception there.
The only expensive part is when the exception does get raised, which costs roughly 4us in .NET and 11us in OpenJDK hotspot (Java has cheaper exceptions when you throw them out of Java, costing about 700ns at throughput, but much more expensive NPEs).
As a result, null-checks are almost always cheap, or free by being implicit (reading array length can be a null-check itself since it will be dereferencing nullptr + 0x08 or so depending on the implementation detail), and the compilers are pretty good at eliding them as well where possible.
If you're using an offset that's not super reliable, and some major platforms of the past 40 years do give you memory at or near zero, and of course with a language like C the compiler is allowed to assume it won't be null and that can cause all sorts of weird program corruption even when a page fault would be safe-ish.
Not on the Amiga ;)
Every time I program something that deletes files I get worried about accidentally deleting the entire filesystem by mistyping something. I shouldn't have to worry about that.
One of the reasons that webapps get as much trust as they do is simply because they don't have unrestricted file access. I wish there was an application format that promised the same on the desktop.
So what I'm saying is, instead of wasting innovation effort focusing on the use of to improve the experience of writing web apps, let's use that to solve the actual problems with web dev.
i think we live in a world where we can do both: we can do wasm and we can continue making improvements to the big three webdev tools.
You could use a more restrictive Typescript subset like https://www.assemblyscript.org/ though.
Also languages like C, C++ or Rust let you exactly define the layout of data on the heap, which is crucial for performance (look up data-oriented-design), and WASM preserves this in-memory layout (since it uses a simple linear heap, like C, C++, Rust, etc... but unlike Javascript, C# or Java). Achieving something similar in a high level language like Javascript would involve mapping all application data into a single ArrayBuffer "pseudo heap", and at that point, it's easier and more maintainable to do the same thing in C (or C++ or Rust).
Having said all that: modern Javascript engines can perform surprisingly well (in general I'm seeing that Javascript performance is underrated, and WASM performance is often overrated - sane Javascript, WASM and native code can all be in the same performance ballpark, but native code usually has the most "optimization potential").
"in general I'm seeing that Javascript performance is underrated, and WASM performance is often overrated - sane Javascript, WASM and native code can all be in the same performance ballpark, but native code usually has the most "optimization potential""
And strong disagree. Javascript is indeed quite fast, but if you use a native compiled wasm libary in the right way (avoiding too many calls to wasm and back) - you will get a worlds difference in performance.
Well yeah, that because each call is basically an "optimization barrier" for the compiler (on both sides of the call), and of course the call itself also adds overhead, although that has been drastically reduced over the years.
Sure, it also is sorta about performance, but JS interpreters are actually shockingly fast in the modern world. Mostly wasm is about moving your existing C++ code (or whatever) for your existing formats or algorithms or tools into a web/mobile client without having to hire someone to recode it all in TypeScript or whatever. That's not sexy, but it's valuable.
1. Desktop: Linux + Windows + OSX
2. Browsers: Chrome + Firefox
3. Phone: Android (no iOS)
4. Embedded
it's absolutely amazing. The biggest problem nowadays is packaging and various runtime things, but what I listed above 100% works, you just need to do the work to package it in N different platforms. How is this not appealing?
Imagine you write Google Sheets, it works on your browser, on OSX, on Linux, on Windows and on Android. It's the same binary.
Personally I think the focus for web technologies should be deprecating old things and being very judicial about adding new things. New things that get added have to be maintained by all browser engines effectively forever, and this is part of why every browser engine is a huge multi-million dollar per year endeavor just to maintain. Meanwhile, a lot of features being added to browsers don't necessarily justify all of this cost. Layout engines are unmanageably complex already. CSS is unmanageably complex already. I think adding even more stuff is hardly the solution, but rather, we need to figure out how to actually utilize what's here better before we can actually come up with successors. Rather than adding more scoped CSS and CSS modules junk, I'd rather just have improvements to CSS-in-JS. Maybe some targeted new APIs that make it easier to implement these features in the browser, not entire new paradigms that require implementing hundreds of thousands of new lines of code that will have to be maintained indefinitely. Likewise for web components: the concept is fine, but every browser has to maintain all of this forever, and all of the edge cases that come from it; will it yield so much benefit from what exists today? Will it actually stop people from shipping megabytes of Javascript, or could they have already stopped doing that if they really wanted to, and all this will do is mop it around a bit? A large node_modules folder disappears when a web app disappears. A large Chromium source code checkout only continues to grow larger effectively forever.
WebAssembly though, gets a pass from me, because it's far more than "Web", but the Web part adds a lot to the overall package. It's just a win all around.
For example, CSS is a disaster. That browsers need to be very complex to implement it correctly is the more reason why it should be replaced.
Or better yet Some Language -> MLIR.
But it turns out many languages do have "Some Language -> WASM" now. WebAssembly brings portability to the table.
No, it's not exactly the same amount of portability as some language and its native toolchain. If you give me source code in some language that I can transpile to C or C++, or into WASM (and then into C), then I can give you a single file which can be executed on Linux, BSD, Mac, or Windows.
That's not remotely the same thing as "its native toolchain."
You, a developer or tech savvy person, might not see any difference.
But if I hand a non-tech-savvy person a file and they can just run it, no matter what OS they're using - with no pre-requisites - that's kind of magic.
Using what libraries & syscalls? A portable IR and a portable executable are very different things, yet the WASM crowd loves to conflate them. WASM is only a portable IR, it doesn't get you portable programs at least not if you want your program to do anything beyond a trivial hello world.
And for Cosmopolitan Libc, there's documented Functions:
https://justine.lol/cosmopolitan/functions.html
And if you want to see things beyond a trivial hello world, you can check out some examples:
https://github.com/shmup/awesome-cosmopolitan
https://github.com/burggraf/awesome-cosmo
Or you can see a pretty big list of pre-compiled Actually Portable Executables here:
Searched for it (probably on Lycos), found one, downloaded the .exe, launched it, popped a CD in. When I pressed play and heard the music start in my headphones, I had the feeling I just performed magic. Big smile on my face as I went back to work.
Viruses, etc. weren't in our reality in those naive and heady days.
I had a text-mode launcher menu for running programs, and that had an "anti-virus" feature built-in that checksummed programs, and alerted you if their checksum ever changed (since these viruses spread by infecting .exes), which is how I found out about it!
Also, in my experience, naive bounds checking will eat up a lot of cycles. But maybe the C compiler can eliminate a bunch of them.
Re: bounds checks, the thing that consumes cycles isn't the bounds check itself, it's Wasm's requirement that OOB accesses produce a deterministic trap, even if the result of an OOB load is never observed and could be optimized out. wasm2c has to prevent the compiler from optimizing out an unobserved OOB load, and that forced liveness defeats some compiler optimizations (probably more than it needs to). But even with all that, we're talking like a <30% slowdown compared with native compilation across the SPECcpu benchmarks.
If you want to transpile arbitrary Wasm to native code in a spec-conforming way, you're probably better-off using wasm2c (which, disclosure, I work on). If you trust the Wasm module, or you're good with the isolation you get from your operating system and don't need Wasm's determinism, w2c2 seems great. Both of these are far less battle-hardened than V8 or wasmtime, especially when you include the fact that now you need an optimizing C compiler in the TCB.
---
* The Wasm testsuite repo has recently merged in the "v4" version of the exception-handling proposal, and WABT is still on "v3". But it does pass all the core tests (including tail calls) at least until GC is merged.
It's much faster to execute than adding a software bounds-check on every load. (Because the module declares its memories explicitly, it's very easy for a runtime to use a zero-cost strategy to enforce that memory loads/stores are all in-bounds.)
But Wasm's safety is more than bounds-checking memory loads/stores. E.g., Wasm indirect function calls are safe, including cross-module function calls for modules compiled separately, because there's a runtime type check (which wasm2c does very efficiently, but not zero-cost).
And, Wasm modules are provably isolated (their only access outside the module is via explicit imports). Whereas if you wanted that from "normal C code," it's a lot harder -- at some point you'll have to scan something (the source? the object file?) to enforce isolation and make sure it's not, e.g., jumping to an arbitrary address or making a random syscall. There's obviously a huge amount of good work on SFI but it's not easy to do either on "normal C code" or on arbitrary x86-64 machine code.
I believe your statement is only true for wasm32 on a 64-bit host where guard pages can be placed around the memory.
Has anyone come up with a zero-cost strategy for wasm64?
This is something that CPU vendors could help with. x86 used to have segment registers but the limit checks were removed in x86_64 so FS/GS cannot be used for this purpose anymore.
(There are probably even better interchange formats coming on the horizon; Zachary Yedidia has some cutting-edge work on "lightweight fault isolation" that will be presented at the upcoming ASPLOS. Earlier talk here: https://youtu.be/AM5fdd6ULF0 . But outside of the research world, it's hard to beat Wasm for this.)
Less important: I don't think going through Wasm has to be viewed as an "extra step" -- every compiler uses an IR, and if you want that IR to easily admit a "safe" lowering (especially one that enforces safety across independently compiled translation units), it will probably look at least a little like Wasm, which is quite minimal in its design. Remember that Wasm evolved from things like PNaCl which is basically LLVM IR, and RLBox/Firefox considered a bunch of other SFI techniques before wasm2c.
E.g. see here:
https://docs.google.com/document/d/17y4kxuHFrVxAiuCP_FFtFA2H...
(I'm not 100% sure but I think that at least V8 has implemented this in the meantime)
My use case would be a wasm based validation lib that could be used in the browser, in edge proxies like cloudflare workers, and in a variety of backends(py/node/.net/java). This approach would negate the need to package a wasm engine like wasmtime.
The WASM GC will be good, not because it's a GC, but because it would allow languages compiled to WASM to interoperate with e.g. the DOM using the same memory manager so you're not e.g. manually managing the memory of DOM objects while JS may hold references to them according to the terms of the GC.
But this also means the DOM API will be tied to the GC, right? I don't know this for fact, but logically that seems correct. So languages that do NOT have good use of the GC are going to be at a disadvantage when it comes to interacting with the DOM, presumably.
Meanwhile, the runtimes of languages like Go and modern Java (after Java gets goroutine-like behavior from project loom) rely on small ~8kb stacks and the ability to swap the stack to a new goroutine using setjmp/longjmp - but WASM isn't a traditional register machine like x86 ASM, instead it's more of a stack machine.. with no setjmp/longjmp.
So not only do languages like Go and Java need to 'decouple' their runtimes for goroutine-like behavior from their GC and make use of the WASM GC instead (which is not tailored to the usage patterns of such languages), but they also need to emulate their own register machine 'on top of' the WASM stack machine so that their goroutine-like behavior works in WASM. Back in 2017 my coworker did this for Go, and so far as I know the implementation has not moved on from that approach since because there is no alternative.
So for Go and Java, you'd be working with a GC that isn't designed for the language and presumably that has performance implications.. and you'd be working with an emulated register machine on top of a sandboxed stack machine (WASM)... that all starts to seem quite a bit far from 'low level like native code' that WASM aims to promise.
I hope I'm wrong, but I fear Go/Java do not have a bright, performant, future when it comes to WASM.
Zig, C++, Rust - all probably have a bright future here. But the challenges for higher level languages are stark
But if you need statically compiled code, no runtime is much simpler to handle, with a lot less to go wrong. But also expect performance that's just "ok". Nothing to call home about.
So what do we lose as part of this round trip?
Just speed? Nothing else?
So, you're doing a classic thing in verification, which is just writing a "checker" that implements some (known good) algorithm in the smallest and most correct way possible, and then using that to check whether much bigger "unknown" things are safe. The goal is that the checker is much easier to implement correctly than auditing everything by hand.
For example, you might have a media player with extensions (like foobar2000); traditionally extensions would be delivered as .dlls, because developers did not release the source. This would be a use case similar to the browser, where WebAssembly would be a good choice instead of random .dll files. They may not want to release the code, but you don't want to trust a random blob of binary code. If you trust your WASM implementation, you don't need to trust that the binary blob is harmless (it will just be rejected or forbidden from doing bad things.)
If you're not dealing with "random binary blob from potentially untrusted source", i.e. you are running the compiler yourself on some code you downloaded, and then running that code, then you don't really need WASM for this, because you could reasonably trust the compiler to uphold the security guarantees using SFI techniques. For example, if you wanted to make sure zlib was safe from buffer overflows in your main process, to reduce blast radius, a pure SFI toolchain would be fine. You can trust it works and then just compile zlib yourself.
But there's generally a lot more mindset and movement around WASM than anything else, so people use it for all of these cases, even cases where they control both the compiler generating code, and where the code is being run.
The main benefit of generating C over LLVM IR is portability: C is supported by far more systems than LLVM can target.
For example, it enables porting Rust applications to Mac OS 9 (https://twitter.com/turbolent/status/1617231570573873152), or porting Python to all sorts of operating systems and CPUs (https://twitter.com/turbolent/status/1621992945745547264).
The main "goal" of w2c2 so far has been allowing to port applications and libraries to as many systems as possible. For more information, see the README of w2c2.
C is a stable format that's backwards compatible for decades; LLVM IR changes with every LLM release. Unnecessarily tying stuff to a LLVM version is a nightmare waiting to happen.
Standard C doesn’t specify any specifically low-level detail, no cache, no vector instructions, nothing.
1. Put more onus for optimization on the converter
2. Mean you can't target as many platforms
So eg, you may generate LLVM IR directly and only get LLVM's optimizations, or you may generate C and compile it with Clang and get Clang's optimizations + LLVM's optimizations
You could always implement your own pre-LLVM optimizations in your LLVM IR generator, but as I think we all know, that's a huge amount of extra work (which is the OP's point)
When you say "70s nonsense" are there specific C features you are concerned about? I would think that transpilation can just avoid bad practices like passing around void * pointers and then casting them optimistically, or even the use of char * for strings in favor of a bounds-checked alternative.
C on the other hand is standardized, stable, ABI agnostic and compiles on pretty much anything that has a CPU.
The hype around WebAssembly is also not very interesting for those of us that have read hundreds of papers on those bytecodes, but hey there is VC money to burn.
In reality, some flaws are preventing adoption. C and Rust are pretty much the only viable languages with WebAssembly support, and since those are also the chosen implementation languages for runtime, nobody is really deriving this benefit. (Anything with a runtime seems to fail for programs beyond "hello world".)
(I tried a "real program" and WebAssembly once, but it ran out of memory when compiled with either gc or tinygo. We had to write the browser side of things in Javascript, which was a shame. The code in question was: https://github.com/pachyderm/pachyderm/blob/master/src/inter.... Somewhat complex, but not so complex you can't do it in Javascript. So a bit of a shame that WebAssembly didn't work out.)
The other thing that I think WebAssembly should fix (but people seem to want to kill me when I mention this) is that the Typescript compiler should just output WebAssembly and we can forget about node_modules and Webpack and that whole nightmare. There is AssemblyScript, but it doesn't run React, so doesn't matter for this use case.
Someday I will go insane and just compile node or deno to WebAssembly and just ship the whole VM with my app embedded in it. Then you'll have a real reason to want to kill me! Muahahaha.
So they are bound to black box attacks, where a clever sequence of function calls can eventually result in something being allowed that wasn't before, as the internal memory state of the WASM module is now corrupt.
Is this an example of WASM not learning from other bytecodes though? 1) Has other bytecodes fixed it? It's not like compiling your C application to CLR protects you from internal memory corruption bugs either, right? And 2) Was protecting against internal memory corruption even a goal? To me it always seemed like the purpose of WASM is to be a cross platform way to run untrusted code sandboxed, just like JS but faster and more appropriate as a compilation target; and whether you compile to WASM or JS, your C program might have memory corruption bugs.
If you're only complaining about the VC-funded hype machine, I could probably be convinced that WASM is sold as something it isn't. I haven't seen it myself really but it's exactly what those sorts of people are wont to do. But that's surely a critique of the hype machine, not of WASM failing to learn from other bytecode designs, no?
https://www.usenix.org/conference/usenixsecurity20/presentat...
One of many security presentations, finding the others is an exercise for the reader that actually cares about security.
There is also no expectation that data in one buffer in webasm would be separate from another buffer.
Show me an exploit, show me with detailed and technical information, show some sort of evidence.
Both of you are just repeating your claim more forcefully, never adding anything that backs up your claims.
Compiling C to webasm doesn't fix C bugs within the local memory space of webasm? No one has those expectations except people with a bizarre vendetta against a simple and benign technology like webasm.
If one wants to argue against pjmlp (as you and I both do), there are two ways to go: 1) you can say that nobody is making those claims, or 2) you can say that if people are making those claims, the problem is with those people, not WASM.
I choose #2, because it seems plausible to me that VC-backed startup types over-hype WASM. If I was more familiar with the discourse around WASM, I might try to argue #1 if I thought I could demonstrate that over-hyping WASM doesn't really happen. Maybe you could go the burden of proof route, that it's their responsibility to prove that people are over-hyping WASM and take the discussion from there.
But you ... don't seem to be engaging with the argument. You agree that "compiling C to wasm doesn't fix C bugs within the local memory space of wasm", yet you say that "the claim of insecurity is not being backed up by anything" when the claim of insecurity is that compiling C to wasm doesn't fix C bugs within the local memory space of wasm. That's literally what was meant by "internal memory corruption" in pjmlp's very first comment which mentioned the security angle. Again, they're wrong to mention this as a flaw of WASM because making C code resistant against "internal memory corruption" was never a goal of WASM, but that doesn't mean that their "claim of insecurity" is wrong.
To show an example of the kind of exploit pjmlp is talking about, just take any C program with some kind of exploitable buffer overflow, run it in wasmtime or wasmer or whatever, and exploit the buffer overflow. Maybe the buffer overflow lets an attacker write 101 bytes to a 100 byte buffer and therefore flip an "isAdministrator" flag from 0 to 1. You, I and pjmlp all agree that such security issues can exist; WASM doesn't magically protect C code from memory corruption. And that is not a flaw in WASM, which is what your argument should be focusing on.
Insecurity would be escaping the VM it runs in. In a native compiled binary you have infinite permissions and can make system calls. You can't do that in a webasm VM.
Webasm is not insecure because a C program that crashes also crashes in webasm.
You still haven't shown any evidence of escaping the VM or actual insecurity.
If someone is dumping privileged data into a VM that's insecure no matter what, why would you blame webasm?
"Why would you blame WASM" is the right question. As I have said in LITERALLY every single comment so far, blaming WASM instead of the alleged hype people is where pjmlp is wrong. He's not wrong in the assertion that insecure programs may remain insecure when run in the WASM sandbox. But you refuse to listen. This conversation is like talking to a much less polite ChatGPT.
What specific marketing are you talking about? Link what you're referring to.
Ask him to link to what he's referring to.
He posts this stuff in half the webasm threads I see. Like you he never has anything to back it up and you both get very upset at people asking.
This is my last response in this thread. You've displayed a profound inability to think. I can't be part of this anymore.
I replied to you because I thought (and still think) you were wrong and/or misunderstanding what pjmlp was saying.
I can disagree with two people at the same time. I don't see the issue.
It is still worth looking at and is actual information, so I appreciate that.
In order to be useful, your wasm application will likely have to be able to make systems calls, or whatever its equivalent might be on your particular host environment. If you can corrupt internal state, you can control the arguments to these calls. The severity of the issue will depend on what your application is allowed to do: If all it has access to is a some virtual file system, the host will still be safe. But if that virtual file system contains sensitive data, results may nevertheless be catastrophic if, say, it can also request resources over http.
given the absolute idiocy some of them finance WASM is a paragon of sensibility and should be injected money via IV. special mentions go out to softbank, tiger global and a16z.
Now we even get kubernetes with WASM containers redoing IBM mainframes/micro computers, Java and .NET application servers as if never done before.
Also, this isn't technical criticism that I asked for.
Who cares what "people are saying" anyway? What is your technical criticism?
I'm not OP, and I don't know if you could have figured this out by just looking at the bytecode formats that came before it, but I think the biggest design mistake of WASM is most likely it being a stack machine. It gives very little (if any?) practical benefits and it just massively complicates everything, both the compilers and the VMs.
I'm not speculating here. I have a pet VM that I'm developing for a register-based IR which achieves roughly the same execution performance of guest programs as wasmtime, but compiles them into native code 160 times faster, doesn't compromise on security, has a bytecode format which takes roughly as much space as WASM, and with an implementation that is vastly less complex (wasmtime's Cranelift is ~150k lines of code; my codegen is, maybe, ~2k lines of code, depending on how you count).
Some of these WASM run-times are literally hundreds of thousands of lines of code running just-in-time compilation. Inferno-level safety hazard.
It takes tens of cycles to instantiate a Wasm module and call one of its exported functions.
There are some serious benefits to OS-mediated hardware isolation, but there are also some real benefits to the "ahead-of-time" isolation you can get from something like Wasm (e.g. via wasm2c->a C compiler->machine code, but also with more mainstream tools like wasmtime).
Launching a new VM is not something that should be done outside of a restart or reconfiguration.
I think for me, what WASM brings to the table is perhaps reduced Linux-isms. Everything has become a little bit Linux-or-nothing, and if WASM presents a unified API towards all operating systems that is a good thing. I'm still not happy that Browsers are de-facto operating systems now, and with WASM even more so.
Link to the talk if interested: https://youtu.be/t4-Al2FoU0k
https://github.com/redpanda-data/redpanda/blob/dev/src/v/was...