(In asm.js, memory was provided by an ArrayBuffer of fixed size, so there memory could truly not grow at runtime.)
55 karma · joined September 4, 2018
(In asm.js, memory was provided by an ArrayBuffer of fixed size, so there memory could truly not grow at runtime.)
I think it actually gives a nuanced answer to your question "where do you draw the line", namely: With a partial precedence scheme, one does not need explicit parentheses for those operator combinations for which the precedence is "clear enough". E.g., in "3 + 2 * 4" precedence is clear to virtually all programmers because it follows standard rules from math ("PEMDAS"), so that is why the precedence table in the post specifies an ordering. However, especially for operators which are less frequently used, or where precedence does differ between languages, one should give parentheses for clarity. I think this is a very sensible argument.
I have certainly made errors because of unclear precedence (e.g., with boolean operators, exponentiation, casting etc.) in the past. And given that there even is a CWE number (https://cwe.mitre.org/data/definitions/783.html) for this kind of error, it seems frequent enough to warrant discussion.
(My understanding of the article is that it mainly argues for _partial_ precedence, not so much that precedence/associativity can be defined in the program - the latter is also the case in other languages, e.g., Haskell.)
You are referring to the example exploit in section 5.3, right? Please note that this example is for a standalone VM, not inside the browser (where JavaScript programs -- and by extension, WebAssembly modules -- do not have direct access to the filesystem).
Whether that exploit is more or less concerning than the browser and Node.js examples, I think is hard to answer in general without additional qualifications. If the standalone VM uses fine-grained capabilities (e.g., libpreopen) or is sandboxed, then changing the file that is being written to might be possible inside WebAssembly memory but access could be blocked by the VM.
However, what we look at in the paper is "binary security", i.e., whether _vulnerable_ WebAssembly binaries can be exploited by malicious _inputs_ themselves. Our paper says: Yes, and in some cases those write primitives are more easily obtainable and more powerful than in native programs. (Example: stack-based buffer overflows can overwrite into the heap; string literals are not truly constant, but can be overwritten.)
As you said, I just hope page protections can still be added later (somebody needs to specify it, embedders need to be able to implement them, toolchains need to pick it up, etc.).
Maybe memory vulnerabilities inside WebAssembly programs can also be mitigated in other ways that do not require such pervasive changes, e.g., by keeping modules small and compiling each "protection domain" (e.g., library) into its own module, or to have a separate memory. I am not sure about the performance cost of such an approach, though.
However, I would draw a bit more attention to the consequences when memory vulnerabilities in a WebAssembly binary are exploited:
(1) Not every WebAssembly binary is running in a browser or in a sandboxed environment. The language is small, cool, and so more people are trying to use it outside of those "safe" environments. E.g., "serverless" cloud computing, smart contract systems (Ethereum 2.0, EOS.IO), Node.js, standalone VMs, and even Wasm Linux kernel modules. With different hosts and sandboxing models, memory vulnerabilities inside a WebAssembly binary can become dangerous.
(2) Even if an attacker can "only" call every function imported into the binary, it depends on those imported functions how powerful the attack can be. Currently, yes, most WebAssembly modules are relatively small and import little "dangerous" functionality from the host. I believe this will change, once people start using it, e.g, for frontend development -- then you have DOM access in WebAssembly, potentially causing XSS. Or access to network functions. Or when the binary is a multi-megabyte program, containing sensitive data inside its memory.
Sure, the warning is early, but I'd rather fix these issues before they become a common problem in the wild.
You are right, some of the issues highlighted in the paper could be solved by compilers targeting WebAssembly. One such mitigation that is (currently) missing are stack canaries. In contrast, stack canaries are typically employed by compilers when targeting native architectures. They also cost performance there (typically single digit percentages), but evidently compiler authors have decided that this cost is worth the added security benefit, since fixing "old C issues" in all legacy code in existence is not realistic.
Note however, that other security issues highlighted in the paper _are_ characteristics of the language, notably linear memory without page protection flags. One consequence of this design is that there is no way of having "truly constant" memory -- everything is always writable. I do think that this is surprising and certainly weaker than virtual memory in native programs, where you _cannot_ overwrite the value of string literals at runtime.
If you don't feel like reading the whole paper, there is also a short (~10 min) video on the conference website [2], where I explain the high-level bits.
[1] https://software-lab.org/people/Daniel_Lehmann.html [2] https://www.usenix.org/conference/usenixsecurity20/presentat...
You are right in that we do not try to attack a concrete host implementation or aim to break out of the sandbox.
I disagree, however, that security inside a WebAssembly binary is irrelevant. If an attacker can arbitrarily read/write in linear memory, it depends on the imported host functions what can be done. The larger WebAssembly applications are (and we believe they will become larger in the future, and incorporate DOM functions, as in our PoC application), the higher the risk that host APIs can be abused. This is troubling especially where there is no sandbox (on some standalone VMs), but even in browsers (where it opens the door for a new type of cross-site scripting).
For example, even if you only call JS eval() with a constant string in your C code, because WebAssembly has no truly constant memory, this can give XSS to an attacker with a memory write primitive inside linear memory.
> The html page there seems to be missing the "main.js" to see what's happening. :(
Oops, good point. They are now pushed, sorry about that.
> main.js, which looks (without checking) like it came from Emscripten
That's correct, you can also check the build script: https://github.com/sola-st/wasm-binary-security/blob/master/...
> C++ compiled to WebAssembly generally manages its own parallel "shadow stack" in linear memory
In the paper, we call the compiler-organized stack in linear memory the "unmanaged stack", to differentiate it from the "evaluation stack" (WebAssembly is a stack-based VM, so this contains arguments and results of instructions) and the "managed call stack" (contains call frames, return addresses, local variables. Managed by the VM, cannot be inspected explicitly by WebAssembly instructions).
> attacker can only jump to a function with a compatible type signature
This is true, but note that WebAssembly types are fairly low-level. That is, there are only four primitive types (i32/i64, f32/f64) and, e.g., a C function that takes a string (char *) and a size_t would be type-compatible with a function that takes a signed int and a struct pointer (all those four types map to the same Wasm type: i32).
Yes, I think several mitigations could be deployed without requiring a change to the language, only by changing compilers and runtime libraries:
* Stack canaries on the unmanaged stack: For storing the reference canary value, you need some "safe" location that cannot be overwritten by regular memory writes. I believe in x86, thread local storage (TLS) / the fs segment register is used. In WebAssembly, one could use a non-mutable global scalar variable, which could only be overwritten by a matching global.set instruction (which cannot be inserted by an attacker). Then, before returning from a function (or really at any point you want to check the integrity of linear memory), you could compare the canary value against this global, as you would in native architectures.
* Allocator hardening, e.g., safe unlinking doesn't require any language support, the allocator just needs to do it (and live with the slightly increased code size cost).
Other mitigations would require language extensions. E.g., for finer grained control-flow integrity (taking source types into account, not just WebAssembly primitive types), one would require multiple table (part of the reference types proposal). Then, only functions with compatible types should be stored in the same indirect call table.
ASLR, unmapped pages, and constants in linear memory are harder, I believe. For ASLR, WebAssembly's 32-bit pointers will most likely provide too little entropy. Unmapped pages and "true constants" would require some host API to make portions of the memory non-writable/readable, but I don't know of any proposal to do so.
> wanting to keep implementation burden low for browsers
> WebAssembly almost a subset of asm.js
Yes, that is also my understanding as to why we are at the current situation. (But I did not take part in WebAssembly's design, so I cannot claim any authority.)
That is good advice. Separately compiling C code into individual modules with small interfaces and as little host imports as possible does make exploitation harder. This is also one of the propositions we make in the mitigations section of the paper.
> If folks start compiling large amounts of their functionality into one single chunk of WASM
However, I also believe that is what will happen soon. Many people are eager to use WebAssembly for more parts of the front-end (and also because they want to develop in Go, Rust, etc., see the increasing number of DOM-libraries for WebAssembly). With more code being linked into a single WebAssembly module, it becomes more important to look into security inside a single WebAssembly module.
I do agree with you, however, that our findings do not invalidate the overall design of WebAssembly. I really like the simplicity and generality of WebAssembly, and the technical (and social) achievement of having a "universal bytecode" supported in four different JavaScript engines (and some standalone VMs), on top of several different hardware architectures (x86, x86-64, ARM, etc.), with good performance, etc. I also think "host security" (which we are not really looking into) is solid, with two qualifications:
* In browsers, WebAssembly is sandboxed like JavaScript is, which is well tested.
* In other host environments (in particular, standalone VMs) the story might be different. It is hard to make any general statements about host security without talking about the concrete host at hand.
> they don't appear to even try to break the WA-host memory barrier
That's correct. However, just because we did not try, I would not say that breaking out of WebAssembly memory is impossible. But those attacks will be engine-specific. There have been some attacks against WebAssembly implementations (https://labs.f-secure.com/assets/BlogFiles/apple-safari-wasm..., https://googleprojectzero.blogspot.com/2018/08/the-problems-...) and surely will be more; a VM implementation is non-trivial with potential bugs after all.
> don't dump WA output you can't validate directly into DOM
Exactly, that is one of the problems with shortened statements such as "WebAssembly has great security". At the very least, one should be aware that inside WebAssembly's linear memory, many common security features are absent. This is relevant in particular when you link together large applications (including legacy C and C++ code) into a single WebAssembly binary, that all share the same linear memory.
Sorry for being a bit late to the discussion, I will try to answer the questions in detail that were asked below.
I want to clarify one misunderstanding that seems to come up several times, namely the distinction between "host security" and "security of the WebAssembly program itself". WebAssembly does have measures and a good design for host security. E.g., in the browser, WebAssembly programs are run in a sandbox (just as JavaScript is), and writes inside WebAssembly's linear memory should never affect values outside of linear memory (e.g., VMs insert bounds checks when reading/writing from/to WebAssembly pointers). Those techniques protect against malicious WebAssembly binaries, which is of course important in the Web.
Here, we look at a different side of WebAssembly's security story: What if the WebAssembly binary is vulnerable and gets fed malicious input? In this attacker model, we can at most do what the host environment allows us to do. But especially for large WebAssembly programs with lots of imports, or WebAssembly binaries for standalone VMs (outside of the browser, without a tried-and-tested sandbox), this can still be a lot of attacker capability! And when we look into the protections inside WebAssembly's linear memory (not between linear memory and host memory), we find that there are very little. All linear memory is always writable, no stack canaries, no guard pages, no ASLR, no safe unlinking in smaller allocators etc. This is worrying as more code gets linked together into a single WebAssembly binary.
If you have any further questions, I am more than happy to answer, here and also via email. Thanks again for the interest!
Regarding your question: No, I have not yet contacted the WebAssembly people at Mozilla. But it's definitely a good idea to talk to the implementation experts. Before I do that, I just wanted to collect more "concrete" questions/problems to ask about.
One of those questions is about the performance overhead of the WebAssembly <-> JavaScript interop. In Wasabi, we have a lot of this, because the "analysis hooks" are written in JavaScript and we insert roughly one hook call per original instruction into the wasm binary. Even without any analysis code, just adding these calls can have a runtime overhead >30x. I would like to optimize this, but before that I need to find our where the overhead is coming from. Possible reasons are (just guesswork, input from people working on this is greatly appreciated):
- that many calls are just inherently expensive, be it cross-language or not (possible solution: be more selective about when to insert calls to analysis hooks) - Wasm <-> JavaScript calls are more expensive than Wasm <-> Wasm ones (possible solution: compile analyses to Wasm, or: hope that this gets optimized better by engines in the future) - the added instructions inhibit some wasm compiler optimization(s) (e.g., inlining is no longer performed because the function bodies are larger than some threshold) - ...
So far, I found working with WebAssembly very pleasing. The spec is compact but still easy to follow. I wrote my own de-/encoder and "high-level representation" of the binary format, which was straightforward and is abled to roundtrip all test files from the spec repo. The most surprising bit was about validation of dead code (i.e., code after an unconditional br is type checked, but the br is assumed to "produce any possible value").
As for personal communication: I am happy to answer any in-depth question via email or so (see http://software-lab.org/people/Daniel_Lehmann.html).
For now, this is in an early stage, so probably mostly interesting to researchers or as inspiration. But I am working on making it more usable (adding documentation, examples).