Not cool.
We were doing this back in the 90s.
The abstract machine behind WASM is a Harvard architecture, so self-modifying code (even accidentally through attacks) is not possible.
Memory safety is still a problem for languages that aren't memory safe, but since the program memory is not writeable from the program itself you can't take over control through memory safety bugs. Now whether the program interprets its data in a way that can be made exploitable is a different bug, but one that exists outside what WASM (or any virtual machine) can fix - it does mitigate the consequences however, since every module is isolated and the runtime uses capability-based security to prevent the code from accessing any system resources that it doesn't need.
The point of wasm is to allow C code to run unmodified, safely, and with as little performance degradation as possible. Sandboxing will mitigate a huge percentage of security vulnerabilities in wasm modules but C is still C. If you want better memory safety, use wasm with a memory safe language.
Wasm downgrades rust memory safety "guarantees" because those guarantees are only truly enforced by the operating system.
Really the only solution is memory safe hardware, which there is effort going into, e.g. ARM CHERI[0] and ARM MTE
[0] https://msrc.microsoft.com/blog/2022/01/an_armful_of_cheris/
For simplicity, pretend any pointer dereference is a function call with the address as its argument. The VM makes sure it’s within the range of valid values every time. There is no executable code in that range either, only data. Further, function pointers are not real pointers, just an index into a table which again, is checked on any indirect function call.
In practice the memory is secured by allocating a 4 GiB aligned memory range and just masking off the pointer on every access so it’s within that memory. This way there’s very little to no overhead, depending on the architecture.
If it is 'use rust' wasm actively downgrades the memory safety that rust gives you because many of the security issues that wasm has is a function of both the compiler and the operating system working together. The compiler merely tells the operating system what should be done - it can't really enforce it.
In this case the various wasm runtimes are acting as that piece and they aren't enforcing the expected behavior.
If the compiler says - hey stick this in .rodata and the wasm interpreter says I don't know what that is then that memory safety feature/contract gets broken.
Besides, that's the safety of the code running inside the VM. Wasm is about the safety of the host.
Can you say more about this? I’m using rust through wasm and I hadn’t heard of any security disadvantages of doing so. What kind of security problems are mitigated in a native binary that are problems in wasm?
By no means an authoritative list and I've not brought up the sandbox at all - that's completely out of scope for this argument but:
* wasm has no aslr - making it incredibly easy to figure out where something lives - you can run the same wasm payload again and again again and get the same result, for instance on linux:
me@box:~/w/zz/aslr$ ./aslr
Library functions: 00007faebfc9f3f0, Heap: 0000562a103c12a0, Stack: 00007ffeea5d8fec, Binary: 0000562a0f1cf179
me@box:~/w/zz/aslr$ ./aslr
Library functions: 00007ff2eda9f3f0, Heap: 000056509ae012a0, Stack: 00007fff849209bc, Binary: 000056509a92d179
~/w/wasi-sdk-20.0+threads/bin/clang --sysroot ~/w/wasi-sdk-20.0+threads/share/wasi-sysroot/ aslr.c -o aslr.wasmand on wasm... Nothing changes:
me@box:~/w/zz/aslr$ wasmer aslr.wasm
Library functions: 0000000000000002, Heap: 0000000000011500, Stack: 00000000000114b4, Binary: 0000000000000001
me@box:~/w/zz/aslr$ wasmer aslr.wasm
Library functions: 0000000000000002, Heap: 0000000000011500, Stack: 00000000000114b4, Binary: 0000000000000001
(btw - all of this works on all wasm runtimes)* wasm can write directly to 0x0 - meaning I can just initialize variables to all sorts of random bad values (eg: overriding functions that might return an expected value or perhaps I have a user_id from a database or worse an 'admin' flag or something, these are just a few examples - it gets way worse in other languages that depend on this)
* wasm can easily overflow buffers (despite claims to the contrary) - that claim really should be taken down from the website - it is totally false
* doesn't have the concept of read-only memory
These are all issues that linux hasn't really had to deal with for 20? years now.
Many systems don't employ ASLR. There's a long history of very smart people questioning if ASLR is even a good deterrent. This dissent even extends to KASLR. ASLR is also not unbreakable, don't fool yourself about this.
Also, ASLR only works if your code is compiled as position-independent, which isn't always the case. There's also (sometimes) a performance hit with it.
Finally, I don't see how the ASLR argument here is anything but moving the goalposts away from your original argument.
> wasm can write directly to 0x0
Again, this is not a guarantee, either. Nothing is stopping an OS from mapping in something at virtual page 0 either. It's just a specification from C (perhaps before) that's been adopted by a lot of other languages, so much so that OS developers tend not to map anything there as a rule.
On many Microcontrollers for example, 0x0 is usually a perfectly valid memory address, and not allowing something there means you lose at least <required alignment> bytes of usable memory.
> wasm can easily overflow buffers
Not sure what you mean by this.
> that claim really should be taken down from the website
Can you quote the claim? Bonus points for linking to it so we have some context.
> doesn't have the concept of read-only memory
So? Again, as many others have pointed out, neither do OSes unless you can somehow guarantee they're mapping memory in with an MMU or MPU with flags lacking X/W. You're still trusting the OS to do it right, the linker to do it right, and the language to place symbols in the right sections.
> These are all issues that linux hasn't really had to deal with for 20? years now.
Of course it has, what are you talking about?
If you don't care/understand about the concept of read-only memory I'm not going to be able to convince you of anything. I'll just state that there's a few tens of thousands of people descending on Vegas next week and the longer the wasm community doesn't deal with these issues the louder those people are going to get.
- WASM has no ASLR.
If I understand correctly, the attack vector here is that if a buffer overrun lets you modify a function pointer, you could replace that function pointer with another pointer to have the program execute different code. As you say, this is hard in native linux programs because of ASLR. You need a pointer to some code thats loaded in memory and you need to know where it is.
In wasm, the "pointer" isn't a pointer at all. indirect_call takes an index into the jump table. Yes, this makes it easier to find other valid function pointers. But wasm also has some advantages here. Unlike in native code, you can't "call" arbitrary locations in memory. And indirect_call is also runtime typechecked. So you can't call functions with an unexpected type signature. Also (I think) the jump table itself can't be edited by the running wasm module. So there's no way to inject code into the module and run it. You only have access to call the already-loaded functions that are in the jump table. And amongst those, you can only substitute one with a matching type signature.
I agree - ASLR could help here. But it still doesn't sound super easy to exploit. That said, it might be much easier to exploit this with wasi when the function table is filled with unix syscalls.
Docs: https://developer.mozilla.org/en-US/docs/WebAssembly/Underst...
- WASM allows writing to 0x0.
You're probably right about this. To be clear, it means if pointers are set to 0 then dereferenced, the program might continue before crashing. And the memory around 0 may be overwritten by an attacker. How bad this is in practice depends on the prevelance of use-after-free bugs (common in C / C++) and what ends up near 0 in memory. In rust, these sort of software bugs seem incredibly rare. And I wouldn't be surprised if wasm compilers for C/C++ start making a memory deadzone here - if they aren't doing that already.
- wasm can easily overflow buffers
Sure, but so can native C code. And unlike native code, wasm can't overflow buffers outside of the data section. So you can't overwrite methods or modify the memory of any other loaded modules. So on net, wasm is still marginally safer than native code here. If you're worried about buffer overflows, use a safer language.
- wasm doesn't have the concept of read-only memory
Interesting! I can see this definitely being useful for system libraries like mmap. This would definitely be nice to have, and it looks like the wasm authors agree with you.
In WASM terms this is called "multiple memories" and it's up to compilers to generate code that works as you'd expect (eg have a "read only data" memory segment and compile the code that references it correctly).
I really don't know how to parse this statement. `.rodata` is not a guarantee, you can almost always change the memory protection of pages your program owns, on almost all OSes (including Windows).
Rust also doesn't make guarantees about this stuff, either. `.rodata` is more borne out of the need to put certain data in certain areas of flash memory depending on the platform. It's also a feature of ELF, not Rust. It's also not even required. The flags specify how the memory is used by the application, not how it's meant to be mapped (though usually the ELF loader will map it exactly how it's specified to be used).
I'm writing an OS. One of the stage 1 bootloader tasks is to load the main kernel as a module. It's in ELF format. I can choose how the ELF segments are mapped into memory, and as long as they're mapped 1) in the right places and 2) with at least the access flags they're specified to use, the program will work correctly.
Further, with Rust in particular, unless a safety contract has been violated within an `unsafe{}` block, I am guaranteed as a contract of the compiler (barring compiler bugs) that anything marked as `const` and thus most likely put into a RO segment will not be written to by the application. That is not a guarantee C can make, and thus not anything an OS can rely on, anyway.
Again, I don't really see your point. I think you're conflating why operating systems have memory protection flags. They're there to help, not to guarantee. You can only guarantee so much, and if you're writing code like what you have in your C example, then the guarantees fly out the window.
I don't see how it's WASMs job to check memory read/write correctness for you. The _whole point_ is to make it as fast as possible whilst being as much of a sandbox as possible - "sandbox" here meaning your application should operate exactly how the language (not WASM) specifies it to work, whereas the environment is locked down to allow for only certain operations and accesses to resources (which has nothing to do about language specifications).
Thus, unless I'm completely missing your point, I disagree with your suggestion that WASM is doing something inherently wrong here.
---
To answer your question directly, barring using something like Rust (which I would recommend people do anyway as it's just a good language regardless of its safety guarantees, but that's a personal and subjective opinion), don't incur undefined behavior in C, and don't rely on the underlying system, if any, to protect you against UB.
main:
.cfi_startproc
.Lfunc_end0:
.size main, .Lfunc_end0-main
.cfi_endprocI was once a huge detractor and naysayer of WASM, but now I'll readily admit that this is the way security "always should have been done".
You can do the same in JavaScript, and it can be very useful. What's your point?
In this case we aren't talking about changing some code in a .js file through a RFI attack. We are talking about changing the behavior of how the node.js interpreter would run the .js file because at the end of the day the interpreter itself is what is getting ran by the os.
If you have specific security concerns, you should be showing actual attacks against wasm runtimes or somehow show that the security model of wasm as a whole will always reduce to an insecure configuration. What you have shown is something that is "exploitable" on pretty much every architecture I know of.
This is another issue:
#include <stdio.h>
#include <stdlib.h>
int main() {
char *s = "world";
s[0] = 'o';
s[1] = 'w';
s[2] = 'n';
s[3] = 'e';
s[4] = 'd';
printf("Hello, %s\n", s);
}How do you propose WASM code is able to change the WASM interpreter behavior?
Could you be misunderstanding WASM's memory model? Are you aware that WASM code can only access explicitly instantiated userspace JS linear memories? Are you aware these memories are just ArrayBuffers in JS userspace?
I'm very far from the only person pointing out a lot of this btw:
https://www.usenix.org/system/files/sec20_slides_lehmann.pdf
https://i.blackhat.com/us-18/Thu-August-9/us-18-Lukasiewicz-...
https://00f.net/2018/11/25/webassembly-doesnt-make-unsafe-la...
You're clearly unable to admit you're wrong. Hard to be charitable about that.