Wasmjit: Kernel Mode WebAssembly Runtime for Linux
github.com
github.com
So, the plausible use is an implicitly trusted app not exposed to untrusted input, that is system-call bound, and that you want to distribute cross-platform without recompiling to targets that have installed this module in a known kernel version. It seems easier just to distribute source for a kernel module; or kernel modules built for the (2? 3 tops) architectures required.
[edit: spelling error.]
The 'system call interface' available to those buggy libraries need be no more capable than existed previously, that system calls are serviced without switching protection mode need not extend the potency of any possible attacks.
A future implementation of this style of design will have much better security and auditability properties than anything that has ever popularly existed in the past. It may have originated with the web, but I welcome wasm and all the market punch it promises to bear on ancient and long-unquestioned corners of our software stack like this. (Naturally I was also a Microsoft Singularity fan)
Or what does WASM do about the concerns that major CPU vendors embed a second co-processor running untrusted, unauditable code with ring -1 access alongside the main CPU?¹
¹which may or may not be active, depending on CPU model and configuration. I've never really fully understood when it is, or isn't, but I think it's somewhat moot.
Further, this assumes the runtime (which is not just the WASM runtime: system calls and other functionality will inevitably need access to the actual, underlying hardware in order to do their job) from being buggy, and since the code is now in ring 0, you get all the consequences of that. (Full, unrestricted access to everything. At least previously, you'd need to find some root exploit to get that.)
(Also, as another poster hints below, since everything is running inside the same protection layer from the CPU's point of view, what prevents you from just running the same attack, in spirit, as Spectre? Only now, there is no protection — everything is ring 0 — so the CPU is "correct" to speculate a read. Sure, the VM will deny it to the code inside the VM, but that didn't matter in the case of Spectre?)
No you don't. Because WASM does not provide many of the facilities that real assembly provides. So, for instance, there is no way to stack-smash using WASM instructions, regardless of the any security problems in the code itself, or even outright malicious code. It just can't do it.
More generally for VMs, there are secure VMs which provide a mathematical proof that the code will, for instance, observe memory safety. Such a proof is much better than ring-level protection, because :
1) You can verify the proof. Good luck verifying the hardware implementation of ring-level security in processors
2) It doesn't take any resources at runtime
3) the proof can be verified at code load time, so insecure code (accidentally insecure or otherwise) just doesn't start executing, ever
4) It doesn't allow for manufacturers to hide "secure coprocessors" or any other bullshit like that
Some runtime stuff implements operations not expressible in JS, and not permitted in the sandbox. You may not want that code in the kernel. Anyway, if random kernel functions can be called from the sandbox, it would be easy to misuse them. Untrusted input might trick code into misusing them from in the sandbox, via nominally valid operations.
Does webasm protect against integer overflow, or unaccounted unsigned wraparound? Not all exploitable bugs are pointer violations.
Processes with a strong user model are already one of the most effective isolation mechanisms in an operating system. Continually hardening your VM/runtime that was ported into the kernel, in my opinion, either results in either you building a microkernel or recreating the userspace/kernel separation that already exists?
While I haven't studied wasmjit, the most obvious implementation is to run system calls naturally, with a global "struct task" existing as before that defines the semantics of the current context, including details like the current UID and capability mask - in other words, without effort, read() and open() could be made to behave identically to before, it's just that the caller now lives in a software sandbox rather than a hardware sandbox, and most/all expensive hardware reconfiguration was avoided
char buff[100];
char func(int idx) {
char *ptr = buff;
return ptr[idx];
}
int main (void) {
printf("%c", func(200));
return 0;
}
Compiles nicely to something like this (module
(type $FUNCSIG$ii (func (param i32) (result i32)))
(import "env" "putchar" (func $putchar (param i32) (result i32)))
(table 0 anyfunc)
(memory $0 1)
(export "memory" (memory $0))
(export "func" (func $func))
(export "main" (func $main))
(func $func (; 1 ;) (param $0 i32) (result i32)
(i32.load8_s
(i32.add
(get_local $0)
(i32.const 16)
)
)
)
(func $main (; 2 ;) (result i32)
(drop
(call $putchar
(i32.load8_s offset=216
(i32.const 0)
)
)
)
(i32.const 0)
)
)
Which will gladly blow up, or not, when i32.load8_s gets called.Where is the WebAssembly implementation that traps on my example?
Relevant documentation http://webassembly.github.io/spec/core/syntax/instructions.h...
Which is my whole point, WebAssembly does not protect memory corruption inside of the module code, which allows for security exploits anyway.
On my sample code if I expose func() to the host, and it gets called with 200 as parameter for a buffer size of 100 bytes, no trap will ocurr.
On a real use case that call might induce an internal memory corruption that will, for example, change the behavior of other functions exposed to the host.
If you wish I can provide an example how to do that, which you can try out in your favorite spec compliant Web Assembly implementation.
Apparently you fail to understand how security exploits are taken.
For example, lets say I have an authentication module provided as WebAssembly, written in C.
The browser makes use of the said WebAssembly module to authenticate the user.
Now we make a cross site scripting attack that calls the WebAssembly functions in a sequence that triggers memory corruption inside of the module, thus influencing how the authentication functions work.
Afterwards the JavaScript functions that call those WebAssembly ones, might authenticate a bad user that would otherwise be denied access.
A contrived scenario that can be easily programmed in https://webassembly.studio/ .
Your criticism is for another layer. If you don't like C, use Rust or Go. They also compile to wasm.
Now we have (at least) three different IR formats that are very similar but have slight differences:
1. LLVM IR. Created for LLVM internal use and may only work with specific LLVM version(s) (correct me if I'm wrong here), is not portable across HW architectures.
2. WebAssembly. Made for web browsers.
3. SPIR-V. Made for OpenCL and Vulkan shaders/kernels.
All of the above have a specification, a binary representation, a text representation and associated tooling. Some (or most?) of the tools seem to operate by doing a source level translation from WASM or SPIR-V into LLVM, then doing optimizations or JIT and possibly converting the resulting LLVM IR back to WASM or SPIR-V. In my experience, the tooling with WASM and SPIR-V is inferior to LLVM IR tools or native code tools (binutils etc).
There's a lot of duplicate work going on here, but I understand that it's difficult to get the stakeholders of Web browsers, GPU APIs and compiler infrastructure to even discuss what could/should be done.
Compiler IR's suitable for AOT compilation (ie. the whole code is compiled before it's executed, although it might be done at runtime) are a game changer on how we deal with programming languages and runtimes and a huge improvement on Java-like bytecode+JIT which never quite delivered what it promised.
What would you use that for, if you had it?
Universities should teach more about mainframe architectures.
Imagine truly cross platform code that performs well, and can be written in whatever language you care for.
The web illustrated certain advantages to developing software that can be easily delivered and run on other systems. However, the learnings of the web were always coupled to the browser.
WebAssembly, through projects like this, may soon be able to empower cross platform code that delivers small, well performing programs on whatever platform a user desires to use.
For user facing applications there remains the question of what sort of cross platform UI framework could accompany WASM. Or should WASM just remain coupled to the Web? Java’s UI solutions were often received very poorly and led to Java’s bad reputation amongst consumer software.
Analyzing the desires and goals of the big players (Google, Microsoft, Apple) with regards to how Wasm might develop is interesting.
Agreed, but want to add wrt tied to Java the language, that stdlib it carries around is a large burden too. WASM will need a stdlib one day or suffer portability, so here's to hoping whichever one is adopted is small and simple (no, posix is not good enough like emscripten and this lib work with).
It doesn't need one stdlib as part of the WASM project, but sooner or later people are going to want to share, e.g. a socket. Even without a full blown stdlib, known data types need to be shared beyond just the structs that are coming with the GC proposal.
If Webassembly wants to remain as agnostic as possible, I don't think there's much more complexity they could add.
Hopefully they get it right, because the JS ecosystem is a bloated mess that we've been stuck with because there's no other game in town.
I wouldn't like running closed binaries with monolithic architectures through my browser. Maybe for a few selected high performance applications.
Not because the security aspect but I think it will be detrimental to an open web. Maybe not, we'll see.
edit: I don't really understand people that are not horrified by the thought of replacing general javascript code with one that is written in c/c++... Because string operations are awesome in these languages, right?!
Plenty of people want it to be a complete replacement for javascript, but I don't see that happening. I certainly don't think the doomsday scenarios that sometimes get mentioned whereby the entire web becomes nothing but closed, DRM ridden WASM binaries is likely. Large corporate owned sites and streaming services will almost certainly leverage it as much as possible, though.
"closed binaries in the browser" have existed since the days of Java applets and Flash. And no one is stopping anyone from providing or distributing the source for their webassembly modules, so I don't see webassembly as being any less free in that regard than any other language that compiles to a binary. Arguably, webassembly is more free because it isn't owned by a company or restricted to a single language.
That it what I was thinking about and it is something I wouldn't want to get back to. And having javascript as compile target would net you a more readable and therefore more open code. Painful process, sure.
But yeah, for specialized application webassembly could very well be worth it.
Javascript and C for example are like completely distinct circles of friends and you don't necessarily want them to meet each other.
I'm not sure where the hate for this comment is coming from though given the submission is also talking about adding a runtime into kernel space.
"Writing Solaris Device Drivers in Java"
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.92....
Notable examples, Burroughs B5500, Xerox PARC workstations, ETHZ Oberon/Modula-2 workstations, UCSD Pascal, ..., iOS bitcode, Android DEX and UWP MSIL.
This isn't that strange of a concept, it's how stuff like Java has worked, but for it to truly feel native the companies that control the platforms would have to coordinate on it and integrate it. Unfortunately I don't see that happening quickly, but perhaps eventually.
The advantage to something like using WebAssembly is you could write a program in C, compile it to WebAssembly, and then it could be run on anything that supports WebAssembly.
This could enable a future where you write one program that can run in a web browser, as a phone app, as a desktop program. All with only tweaking the UI.
Additionally the divide between a “web app” and a desktop app from a user perspective increasingly comes down to: do I need this to run fast (as a native program), or do I want it to be convenient to use (as a web app)? WebAssembly may very well become fast enough that it is fast and convenient for all but the absolute most demanding of applications.
The C ABI describes a common calling convention and a few other details to allow different compilers to use each other's binaries, but C is hardly a common language in the way WebAssembly is meant to be.
https://slikts.github.io/js-equality-game/
It is all bound to make JS kernel programing a Pilar of OS Stability.
It's weird nobody seems to really complain about platform lock in, which is the main problem created by software companies, that create so much pain for developers.
Even if one looks at the software market, I really doubts that lock in really works at all. Developers quickly figure out ways to do things run on multiplatform.
It's no surprise JavaScript grew to be so popular, because it was just 100% multi platform. It's quite sad to realize js has so many drawbacks.
For companies like Microsoft and Apple those delays have been enormously profitable, but long term there may be a hefty cost from their avoidance of truly great crossplatform software development tools.
If kernel module developers become a vocal part of the wasm ecosystem, I'd expect 8- and 16- bit data types to get added before long.
They are handling endian issues by simply requiring little-endian.
It makes one wonder what would have happened if the Micro/370 had gone forward and the architecture-neutral userspace had evolved in the PC space.
http://www.cpushack.com/2013/03/22/cpu-of-the-day-ibm-micro-...
The paranoid security engineer in me screams "Ring0 is what I always wanted my browser to have access to".
Interesting project!
Just, erm, load these binaries provided by FB, Google, Microsoft and Netflix. :P
Just use MMU hardware to insert 2GB [0] no man's land below and above Wasm memory and only allow signed 32-bit indexing (-2^31 — 2^31-1).
This way the attacker can only read sandbox memory (and the useless 2 GB no man's land, mapped to pages full of zeroes or whatever).
When there's no speculation involved or the data is simply out of "speculative range", Spectre is toothless.
[0]: 2GB is just a basic example. More may be required if for example something like x86 SIB (Scale Index Base) is used for multiplying the index by 2, 4 or 8.
That's a good idea, but there's also several reasons that may not be appropriate:
- WASM explicitly says that it may be extended to 64-bit indexing (more than 4GB of addressable memory is definitely useful for some things)
- Spending 4GB of (hopefully, virtual) memory on every WASM instance may be undesirable or impossible (e.g. 32-bit processor)
That said, it's very reasonable to impose restrictions on things running in ring-0, and wasmjit could well require a 64-bit machine with 32-bit WASM indices (which I imagine would be okay assumptions for things one would do with it anyway).
In that case, just fall back to bitwise AND index clamping. A small performance penalty, but nothing major.
> - Spending 4GB of (hopefully, virtual) memory on every WASM instance may be undesirable or impossible (e.g. 32-bit processor)
Just page table entries. Wasting physical memory for that would be pointless. If the entries need to be mapped, on x86-64 it'd incur 4 kB, 2 MB or 1 GB total "wasted" memory, depending on which page size granularity you want to use. Of course, you could also simultaneously use this "wasted" memory for any non-sensitive data.
Well, mapping 2x 2GB memory using 4kB pages does take up hmm... 8 MB of RAM for the PTEs. So perhaps 2 MB pages would be optimal.
Masking the index will break code that is actually using the larger address space: running true 64-bit WASM code (as in, using >4GB of space) won't work, which is what I was referring to.
> page table entries
Indeed, hence the reference to virtual memory. In any case, because both x86-64 and ARM64 only have 48 bits of actually addressable space, that 4GB of overhead (plus, up to 4GB of actual addressable memory) only allows for 65536 (or half that) WASM instances. That's definitely a large number, but not one that is out of reach.
You can also clamp for example at 33-37 bits, giving 8-128 GB array range.
cmp %bound, %ptr
jae bad_ptr
sbb %mask, %mask
and %mask, %ptr
Just two extra instructions. No need to memory map or hard code the size of bounds.See `array_index_mask_nospec` in https://github.com/torvalds/linux/blob/master/arch/x86/inclu...
Pretty neat idea! [Although the (register) dependency chain looks a bit nasty. 'and' will need 'sbb' to commit and 'sbb' will need to wait for 'cmp' to commit (flags register). But I guess the few/rare cases where this latency is really an issue can be dealt one-by-one basis.]
> No need to memory map
Well, using MMU can have performance benefits. Less repetitive bounds checking code and better performance in most scenarios. Both solutions have their strengths and issues, there are no silver bullets.
You don’t need to use one exploit if two that interact are common enough.
How about a nice pointer arithmetic bug in any reasonably common version of the webasm runtime?
So you can still trigger memory corruption if the WebAssembly was generated from languages like C and C++.
It's indeed possible to avoid ever generating speculative fetches to sensitive data by simply ensuring the indexed access can't ever reach outside the sandbox in the first place, not even in any possible speculated case.
Just don't use conditionals and branches to do that, but something else. Like MMU to move data out of range or bitwise AND clamping
There doesn't yet seem to be a implementable standard to define what a non-browser execution environment looks like. The beginnings of a common spec are here though, which some places have started working with:
https://github.com/CommonWA/cwa-spec
That'll likely need to make it's way into the official specs:
* https://github.com/WebAssembly/design
* https://github.com/WebAssembly/spec
At which point implementations will have something to focus on. :)
For the Go, this is the (recent) matching issue:
https://github.com/golang/go/issues/27766
Kind of guessing it'll turn into a tracking issue to get it done.
JS already unhinged from the browser pretty thoroughly. I think when it comes to Wasm it's almost as much about what it doesn't have as what it does have. Lack of DOM bindings and a GC make it much more suitable for hosting in more environments like the kernel.
Reminds me of "Why we use the kernel's TCP stack"? And eternal debate over the performance benefits of "kernel bypass" and "zero-copy" technologies such as DPDK versus full userspace TCP implementations like OpenOnload.
https://blog.cloudflare.com/why-we-use-the-linux-kernels-tcp...
Conclusion is that the untapped fruit in kernel TCP gains comes from tuning performance of the network stack itself: optimizing overhead a packet incurs on its path through the stack. A very active area of development.
Netdev 0x13, THE Technical Conference on Linux Networking
[0] A (half) joke from Gary Bernhardt's excellent The Birth and Death of JavaScript (pronounced "yavascript"): https://www.destroyallsoftware.com/talks/the-birth-and-death...
That said, this claims it relies on emscripten, I think? You can do that with Rust, though it’s not the preferred way. I sure about Go.
Emscripten provides a libc; I wondered if it required it. I guess not.
I think you may want to reassess your definition of what a runtime is ;)
> "crt" stands for "C runtime", and the zero stands for "the very beginning".
[ https://en.wikipedia.org/wiki/Crt0 | https://en.wikipedia.org/wiki/Runtime_library ]
Yup. C definitely has a runtime!
Wasm, on the other hand, is a memory-safe interpreter, but unlike eBPR is Turing-complete, so it can run arbitrary programs.
Well, almost.
> The main object of this patch series is to support verification of eBPF programs with bounded loops. Only bounds derived from JLT (unsigned <) with the taken branch in the loop are supported, but it should be clear how others could be added.