[0]: https://www.destroyallsoftware.com/talks/the-birth-and-death...
[0]: https://www.destroyallsoftware.com/talks/the-birth-and-death...
In the video, his argument was that the browsers are single-process anyway, and if everything runs in that process, we don't need that separation. However, since then, we've learned that single-process browsers are a security nightmare, so these days browsers are actually not single-process anymore to provide proper sandboxing.
But I love how close to correct that video is, and it's interesting to see in what ways it turned out to be wrong.
Apple has spent a long time hardening the JavaScriptCore web sandbox to run untrusted code. We’ve come a long way since JailbreakMe’s web-based jailbreak, but ultimately memory safety requires participation from all parts of the stack and JavaScriptCore and V8 are still both written in C++. You can trigger memory-safety vulnerabilities in the host VM using guest code.
wasmtime is supposedly a hardened WebAssembly runtime written in Rust, but it’s also a JIT, and I have no idea if anyone has put it through its paces security-wise yet. The idea is that WebAssembly can have JIT-like performance without JIT-like security concerns thanks to a simpler translation layer and minimal runtime.
I could see an argument for dropping some layers if the VM isolation become stronger
No its not, when it comes to end-user app performance, experience or privacy.
Sure, by adding security we can have another reason to let developers end up with golang app compiled to wasm running within electron sandboxed through API redirection (OS + antimalware/antivirus/BPF based EDR) and use it for, like, listening music in a very secure way..
With all these layers happily streaming all kinds of telemetry to knows where, with owning nothing but a bunch of numbers behind a ton of DRM layers, and with no ability to change things to the point where we can't have an app's theme matching system colors because crossplatform compatibility/security reasons.
Case 1, firefox:
> dom.security.unexpected_system_load_telemetry_enabled > security.app_menu.recordEventTelemetry > security.protectionspopup.recordEventTelemetry > security.certerrors.recordEventTelemetry
I don't want to accept developer's assumption that these have to be enabled by default.
Case 2, Windows: can't even do a build of a trusted codebase under IntelliJ without antimalware adding up, like, +150% to build time. While IntelliJ (or some of its extensions or plugins that creep up during development) is happily reporting that performance issue back to its masters. Ugly.
I can't put my finger on it, but somehow this looks familiar (hint: it starts with a "J", too)
So I don't think there's an automatic win in terms of performance by ridding yourself of it. Especially if you're running through the (pretty slow) WASM VM layer anyways.
For some applications (e.g. databases), running unikernel or closer to kernel and having direct access to the MMU could be a big win (e.g. see https://github.com/tuhhosg/exmap & https://github.com/viktorleis/vmcache & https://www.cs.cit.tum.de/fileadmin/w00cfj/dis/_my_direct_up...).
For general applications esp those written to a POSIX standard or making assumptions that the machine they're running on looks like a typical modern day computer? Dubious. You'd end up writing a bunch of what the VMM layer does in user code.
the more I think about it the less it makes sense
- js engine rely on vmm, and wasm does so, too (in many ways)
- close to every non embedding, non trivial program I have seen is in subtle ways based on the assumption of vmm
- some vm technology, especially around micro vms uses vmm, too. And Unikernels only really make sense as VMs
Could you expand on this?
The “javascript instruction” is FJCVTZS, which is a rounding mode matching x86 semantics, which is incidentally what JS specifies for double -> int32 conversions, and soft-coding it on top of FCVTZS it is rather expensive (it requires a dozen additional instructions to fix up edge cases).
This is beneficial to javascript (on the order of a percentage point on some benchmarks suites, however pure javascript crypto can get high double digits gains), but it’s also beneficial for any replication of x86 rounding on ARM, including but not limited to emulating x86 on arm (aka Rosetta 2).
The hardware mmu does have costs: tlbs are quite small and looking things up in a several-layer tree adds a lot of latency. If vm were fine, no one would care much about hugepages, and yet people do care about them. (Larger pages means fewer tlb misses and fewer levels in the tree to look up when there is a miss)
> Consequently, modern processors have extremely large and highly associative two-level TLBs per CPU — for example, Intel’s Skylake chip uses 64-entry level-1 (L1) TLBs and 12-way, 1,536-entry level-2 (L2) TLBs. These structures require almost as much area as L1 caches today, and can consume as much as 10 to 15 percent of the chip energy.
Bhattacharjee, Abhishek. "Preserving virtual memory by mitigating the address translation wall." IEEE Micro 37.5 (2017): 6-10.
However, the 4kiB page size that is typically used and is baked into most software was decided on in the mid-1980s, and is tiny compared to today's memory and application working set sizes, causing TLB thrashing, often rendering the TLB solution ineffective.
Whatever overhead software memory protection would add is likely going to be small in comparison to cost of TLB thrashing. Fortunately, TLB thrashing can be reduced/avoided by switching to larger page sizes, as well as the use of sequential access rather than random access algorithms.
As I understand, the Smalltalk world put a lot of engineering effort into making the software-based model work with performance and efficiency. I don’t think the results were encouraging.
A lot has happened since Smalltalk.
Then again, that's the theory, in practice there are many reasons why hardware moved from early segment based architectures to paging, and memory isolation is only one of them.
segmentation was an evil everyone both from the hardware and software side was very happy to get ride of
whoever reintroduced segmentation will probably be burned on a stick by computer developers in the afterlife (/j)
Then, at runtime, it then does nothing at all. Which is very fast.
[0] https://www.microsoft.com/en-us/research/publication/deconst...
Also spectre.
You could of course add dedicated hardware to lower the overhead specifically of memory access permission checks. In fact most CPUs already do, it is called an MMU.
[1] https://www.theseus-os.com/Theseus/book/design/idea.html
Maybe watch the project founder's talk? https://youtu.be/n7r8zO7SodE?si=nswWcFrkTj7K1GpZ
Swapping object graphs out to disk (and substituting entry points by swap-in proxies) was a thing in Smalltalk systems, and I expect Lisp machines must have had their own solutions. For that matter, 16-bit Windows could (with great difficulty) swap on an 8086, and other DOS “overlay managers” existed. Not that I like the idea, necessarily, but this one problem is not unsolvable.
I still remember using overlays on Turbo Pascal, Turbo Basic and Clipper.
Amiga also didn't had a MMU, and we all "enjoyed" our Guru Meditation momments.
An exploit to the runtime in such a system obviously would of course be a disaster of upmost proportions, and to have any chance of a decent performance you'd need a very complex (read exploitable) runtime.
The question is.. if you have full isolation and separation of the processes etc... why are you bothering with the WASM now?
WASM can help with portability.
Any sandbox layer can help with anomaly/exploit/bug detection, accelerating fixes to untrusted code, or a neighboring sandbox layer.
"Phrack: Twenty years of Escaping the Java Sandbox" (2018), https://www.exploit-db.com/papers/45517
> Put some WASM in a JVM in the WASM. In an OS. In a hypervisor.
Intel TDX comes to mind.
One more recent effort that also implements the same idea is the Phantom OS [1].
In reality, I think there is always going to be a hypervisor to separate the various workloads, and the hypervisor is likely to keep using paging, to support dynamic memory partitioning -- though perhaps with a larger page size, so as to not create too much pressure on the TLB.
The CPUs were microcoded, so the bytecode was for all practical purposes Assembly.