WebAssembly: Docker Without Containers
wasmlabs.dev
wasmlabs.dev
I see several languages moving towards more and more WASM but on a technical level I don't see the benefit of WASM over something like Firecracker. Docker and other sandboxes have to deal with shared kernels and all the risks associated with that, but leveraging virtual machines instead solves that issue. There are already proof of concept implementations to replace Docker with VMs as a virtualisation layer, so I wonder if it wouldn't be better to invest time in getting those wrappers completely up and running rather than coming up with essentially "Java but we also emulate the OS".
Until WASM advocates start including benchmarks in their blogs, I'll keep watching this stuff from a distance.
Knowing nothing about WebAssembly, I would guess it's because JS runs on a single thread.
https://web.dev/workers-overview/#:~:text=Web%20workers%20an....
https://developer.mozilla.org/en-US/docs/Web/API/Web_Workers...
https://developer.mozilla.org/en-US/docs/Web/API/Service_Wor...
Whether or not workers are actually implemented as threads or processes in the runtime is irrelevant. As far as the JS code itself is concerned & what you can do with it, browser JS is lacking multi-threading. There's just no way to do a shared heap, and that is the biggest defining difference between a process and a thread.
Applications like Photoshop and Google Earth use ptheads on the Web so their compiled C++ is multithreaded, very similar to how it would run natively, and with similar responsiveness and throughput speedups. Though there are some limitations too, see
Should address that concern. However there is another way of looking at performance and is in the context of serverless where typically single threaded performance is inportant, as well as cold start time etc and that’s why Wasm is popular in that scenario
WebWorkers don't give you multi-threading behaviors (heaps/address spaces are not shared). WebWorkers would be how you launch a new process, but there's still otherwise no way to make a thread (nor even a fork() equivalent for that matter).
I could be wrong here, I haven’t done it, but I thought this was the reason for supporting atomics in the first place.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
As in, SharedArrayBuffer is equivalent to shm_open. Which means it's not even that good as a shared memory construct as it's missing all the protection enforcement of memfd (or Android's ashmem)
And no, WASM doesn't support memory protection
Look at old-good JVM. It has tons of tools to analyze and understand behavior of your production system. You could have thread dumps (stack traces of all existing threads) at any moment with negligible performance impact, you could dump heap and analyze it off-site, you could have tons of metrics, about each dark corner of mutexes, GC process, about JIT, including, if you need it, generated native code!
Many of these thing you could get on production, not in sand-box.
If you system behaves strangely, live-locks, consume more memory than you think it should, tharsh GC, you name it, you have all tools to understand what is wrong, find bugs or mis-configurations, etc.
With all these new-and-shiny WASM and not-so-shiny JS VMs you mostly in the dark now. Service become unresponsive? latency goes to the roof? Only thing you could do - restart.
It is not property of WASM per se, but this infrastructure is too immature now, comparing to 25+ year old technology.
But from my experience most people don't know these tools even exist so the only thing they do is restarting and guessing where the problem might be if it persists.
Any JVM anywhere can answer the question "why am I running slow" with a quick run of flight recorder. Memory, CPU, socket time, GC impact, TLB, thread dumps, etc. It's all there in one file that imposes something like a 1->2% performance impact if you run it constantly.
It's just so good.
but you can have threaded wasm code in any evergreen browser since a more then a year as far as I'm aware
Basically the trick is that you use multiple web-workers with the same WAS program and the same shared buffer. Then you also add some JS glue code to coordinate which thread is the main thread and which threads you use as thread pool (e.g. in rust/wasm with rayon you can set it up as worker pool).
Now there are some drawbacks (last time when I used it, might have gotten better):
- threads are started/managed from outside (so don't expect any kind of "spawn" function to work, generally spawning new threads is non-trivial and so is (properly) cleaning up old threads, through if you need a fixed worker pool it's all fine)
- there where some limitations wrt. threading/synchronization which made certain usages of concurrency rather slow (through many where fine)
- no "synchronized" operations mustn't be called from WASM code called by the main JS thread. This means in most situations you need to pass data to web workers and then to WASM (instead of e.g. passing it to WASM and then using in wasm a mpmc-channel to pass it to the worker pool). There are some optimizations around passing pointers as numbers to/from the web-workers but it's limited and not nice. Or at least wasn't ~a year ago.
- bugs in Safari leading to strange crashed for code running in all other browsers nicely under unclear and non-debuggable circumstances (probably fixed, I hope)
Anyway all in all using rust->wasm with rayon and a thread pool was already surprisingly viable ~1 year ago.
What do you mean? This bullet point had a rocket emoji! Surely you don't actually want evidence to support a rocket emoji?!?
https://www.fastly.com/blog/lucet-performance-and-lifecycle https://arxiv.org/abs/2010.07115
With Firecracker I believe snapshot restore time is around 2-3ms. In my tests wasmtime ran about 50% the speed of native so depending on your workload it might still be faster for short running jobs where the startup time dominates. (Wasmer was maybe 80-90% of native speed but I don't know their startup times.)
WASM startup isn't going to be any faster than native code startup. It's going to be strictly worse if anything thanks to the JIT, although you can AOT that to native and then just restore parity with native code.
Which just gets back to the speed of your startup depends on what your startup does.
Some of those techniques can be applied to native code too, see "On-demand-fork: A Microsecond Fork for Memory-Intensive and Latency-Sensitive Applications" https://www.cs.purdue.edu/homes/pfonseca/papers/eurosys21-od...
But I think wasmtime can always be faster to instantiate since the guarantees provided by the runtime allow it to safely reset and reuse instantiations:
"We implemented an “instance allocator” in Wasmtime that makes use of this copy-on-write (CoW) technique for very fast instantiations. It also uses a Linux syscall known as madvise to quickly “reset” the page mappings back to the original read-only heap image, so we can reuse the same mappings over and over when the same Wasm program is re-instantiated many times. (One might imagine this would be the case in a server serving many requests, for example!)"
https://bytecodealliance.org/articles/wasmtime-10-performanc...
I doubt this will ever be as fast
But it also depends a bit on the application.
Some applications can benefit a a lot from CPU specific instructions combinations which are not available to wasm (with available I mean implicitly, i.e. your wasm code gets compiled to them).
Luckily for a lot of use-cases this doesn't matter much(1) and some degree of SIMD support is often(2) available.
(1): Without micro-optimizations which most times aren't done as due to their maintenance/development cost.
(2): I'm not quite up to date. I think 128bit SIMD is available in most (all?) relevant WASI runtimes and at least some browsers.
when I mention performance, they kinda waffle a bit, saying "CPU is cheap" or something similar, and they start to show a hint of understanding when I say that cloud resources are billed by unit of CPU time, and by amount of RAM used. then I say that our mutual employer invokes lambdas hundreds of trillions of times per year and I think they briefly understand before being caught up in "new stuff is awesome" technology fetishism again.
it's exhausting.
everyone should live overseas for a couple years because it changes how you view the world... everyone should be a game developer for a couple years as well, because you will quickly notice just how unbelievably slow modern software is. more people need to see that.
security is important! portability is important! other things are important, always, and when you gain a sense of just how slow software is today in comparison to how unbelievably fast modern hardware is, it becomes very hard for me to think positively of anything that lowers performance further for almost any reason.
This is an easy engineering problem which will be solved when there's enough motivation and engineers working on it.
Increasing adoption is more of a business problem though, and it is unclear if performance is the bottleneck here.
If you have a wasm product that has to be as fast as native code, the solution is to find compiler engineers (or a company that specializes in this) who will solve this for your situation.
You can do a lot more if you just want to speed up your codebase.
All the big tech companies employ multiple hundred compiler engineers each for this purpose.
- Solomon Hykes (co-founder of Docker)
That is what this whole trend is all about, replicating application servers with WASM.
Every time I need to dive into k8s stuff, I can only think "this was so much easier when configuring WebSphere and EAR deployments".
But the similarities are shallow at best and tied you into a single VM.
Saying “this whole trend is all about replicating application servers with WASM” makes it sound derivative. I mean, yes, but only yes in the same way “cloud computing is just replicating large remote shared mainframes”.
And applications get tied to the WebAssembly ecosystem, it is also a single one.
For one, the idea is that it should be fairly simple to compile an arbitrary program to WASM, allowing you to use a far wider variety of languages.
In this case, it’s more akin to “docker with extra steps” as opposed to “docker but you can only hire Java devs”
Wasm is a refinement of the ideas in the JVM, with a couple good JVM on Wasm solutions already existing. The best of which is CheerpJ.
Instead of referencing NestedVM, you should link to GraalVM which I know you are aware of.
The main difference with regard to compiling “arbitrary” programs between the JVM and WASM is that the JVM doesn’t have untyped linear memory like WASM does. WASM isn’t type-safe within the linear-memory regions (which is what allows C-like languages to be compiled more directly), whereas JVM ensures type safety for all objects.
Saying that “WASM is a refinement of the ideas in the JVM” is simply wrong, they have different design goals and therefore implement different design trade-offs.
As you say, it can be trivially emulated, no reason that linear memory couldn't be implemented on the JVM using an array of primitive (char,int,long).
The Wasm VM is more generic and has a better capabilities security model than the JVM. Had the JVM been more like Wasm (signed and unsigned primitives, capabilities model), it would have been a natural compilation target for sandboxing native code.
As Wasm gets more features it becomes more like the JVM (GC, reference types, component model). Eventually Wasm will subsume all the features that differentiates them.
Why is that different?
WASM is pretty different to the JVM. The JVM deals with a lot of higher level constructs like objects, constructors, virtual methods, the GC etc. Which is fine if the language you’re hosting works in that way.
WASM is more like assembly - the raw intrinsics used by a hypothetical WASM CPU. So you can compile a lot more stuff to it, because it’s a more natural target than a much, much higher level VM like the JVM.
On the browser side, you could maybe make a ChromeOS style argument of just wanting to run everything in a browser no matter what, it's nice for it to be sandboxed etc. But if you look at what your Postgres is doing it's sort of a Linux VM, and you can run those already at full speed: that's WSL and there are various solutions on macOS like Docker Desktop.
On the server side, the reason nobody bothered trying to run Postgres on the JVM before Graal is that it's unclear what the benefit is. Security? Processes and kernel isolation seems to work well enough on the server, and anyway Postgres is a highly trusted component anyway. CPU independence? That hasn't been useful since stuff like SPARC died, we are now just starting to see ARM chips appear on the server, but compiling native code for both intel and arm isn't hard and Linux distros have the infrastructure to do it for a long time already. For custom servers, well, not many are writing them in C/C++/Rust anyway these days. RISC-V remains mostly theoretical.
To what extent is this doing it because it can be done, vs delivering real benefits? My mind is open but the low level nature of the WASM instruction set also means the VM can't offer many benefits to the programs inside it, nor the users outside it.
The JVM has historically and famously sucked at sandboxing untrusted/partially trusted code. The JVM also isn’t a suitable compilation target for arbitrary and existing codebases.
WASM is built to sandbox untrusted/partially trusted code. WASM is built as a compilation target rather than a complete hosted VM with bells and whistles.
There are advantages and disadvantages to this. One advantage is language-agnosticism. Another might be around the ability to run user-supplied untrusted code in a much safer way. See Cloudflare functions.
One disadvantage is the lack of bells and whistles.
So, the answer to “why is this different to the JVM” is that.
We need to drill into this a bit more, because the WASM ecosystem can certainly learn lessons and do better than the JVM but this isn't quite the right set of lessons to learn.
The JVM spec was written from day one to sandbox arbitrary and partially trusted code. The SecurityManager architecture is now being removed, but that's a big project exactly because it was deeply integrated into everything. It's also embeddable and a compilation target (it interprets bytecode), and later it was extended with features designed to make it more language agnostic as well at least for everything at the same level of dynamism of Java or higher (so indeed not C/C++ but yes for most other langs).
So the interesting thing to debate here is not really the design goals, which are very similar, but what concretely will make WASM more successful at achieving these goals.
For example, what made the SecurityManager difficult wasn't something fundamental to the JVM but rather that code which is both useful and sandboxed needs APIs that let it do privileged things in controlled ways, and a lot of the bugs were in those API implementations or on the boundaries. That's especially the case on the desktop where sandboxed code had to call into a lot of OS libraries.
On the web this problem is solved with the browser makers exposing JS/renderers to WASM or just saying do the tricky stuff in JS, and then wrapping the whole thing with kernel sandboxes and IPC to handle the fact that the JS/WASM sandbox itself will inevitably fail in the same way. In other contexts this, well, doesn't seem to be solved, really? The moment you start exposing APIs to the WASM code you face the same problem. Also these days you have spectre to think about, so maybe you need a separate process sandbox anyway. Alternatively you can do what the GraalVM guys are doing (in their EE) and using Intel MPKs but that's pretty advanced and I didn't hear about anyone else doing that.
Now there are still some crucial differences! The JVM wanted to allow sandboxed code to interop smoothly with higher privileged code. This opened up a bunch of reflection-based bugs whereby you could reflect your way to the SecurityManager and switch it off, confused deputy attacks and so on. There's no equivalent in WASM, but that's partly because (current?) WASM doesn't really try to define a fine grained permissions model or a way to mark some bits of code in a program as more privileged than others. It also doesn't provide a large set of pre-implemented APIs. The sandbox is whatever the developer exposes to the context. This doesn't make WASM stronger, it just means it punts the really hard bits to the user i.e. browser devs. GraalVM's new JVM sandbox (not SecurityManager based) works mostly the same way, as do process sandboxes so this is definitely the trend, but of course there was a reason the SecurityManager was created that way and it's because it requires way more code and work by the developer to sandbox code if you don't have support for tight mixing. So maybe the sandboxes that do exist will be stronger, but there'll be less sandboxing overall and permission scopes will be much wider. Is that the right tradeoff? I'm not totally sure it is but eh, people really like all-or-nothing and that's the way the industry is heading now.
At any rate that discussion is a bit academic, because you can't do an OOP capabilities type architecture in C anyway.
What about language agnosticism? Again the hard part here isn't having a common bytecode - CPUs already provide that - it's all the engine bindings and semantic alignment required. If you want to pass a std::time into JavaScript then something has to bridge that gap, if you want to call into a dynamically typed language from a statically typed language, then something has to generate interfaces for the compiler to check against and so on. Here I don't see what WASM has to do with anything really, it's just not in scope. Compiling a Python interpreter to WASM doesn't make it any easier to call Python from C++, or Java, or JavaScript. The SOTA there is Truffle/JVM by far.
Wasm can learn from the JVM for sure, but if you went almost 30 years back in time, you would have some great tricks to teach the JVM, learned from Wasm.
Consider .NET / CLR for another example. Like JVM, it also deals with objects, methods etc conceptually on bytecode level. But it also deals with raw pointers and pointer arithmetic and other such stuff. As a result, you can efficiently compile e.g. C into that bytecode. So wasm isn't really new in that sense, either.
You can run Postgres using WASM in two different ways[1][2]. That’s a non-trivial codebase.
1. https://www.crunchydata.com/blog/learn-postgres-at-the-playg...
You can compile arbitrary programs to WASM, like you can compile arbitrary programs to x86 or ARM. It’s effectively a CPU target. You can take the whole of Python and SQLite, compile them to WASM and run a web framework on top of that via WASM in the browser[1].
If those programs utilise specific os-level behaviour that can’t be shimmed, or explicitly throw an error when being compiled to WASM then of course they won’t work without modifications.
Meanwhile, the JVM is not anything like a CPU target. It’s a very high-level VM designed for a particular type of gc’d and jitted language.
If the model works, why shouldn't there be competitors and wide spread adoption?
WASM likely won’t convince all languages to switch to its bytecode by default. Its adoption story requires enough people to maintain this non-standard compilation pipeline.
Which is fine, but beyond the bytecode, the productivity will only be maintained if it offers an equal or superior interface to POSIX. Right now, WASI is not it. A lot of basic elements are experimental, like file seek, multi-process, threads, SIMD, GPU programming…
People will believe the hype, miss the caveats, use it at work, and have their project fail. Companies will blacklist the technology.
In my mind, it is too early for them to be so publicly dithyrambic. All WASM publications should link to a page that details all WASI features that are still experimental.
.NET/Java/Clojure, maybe not. Ruby and Python work via an interpeter for the platform, just like they do on JVM/.Net, except that, unlike JVM/.NET, in both cases they can use the normal C-based interpreter, compiled for WASM.
There is no world where people are just grabbing an existing app and saying "hey, I'm gonna drop this into my wasm runtime real quick"
This is great for running existing "native" type code compiled from C/C++/Rust but its profoundly unsuited to the kind of development that most application or service developers do, which is in higher level languages with automatic memory management, monitoring / profiling services, etc. All of which either have to be re-invented in the WASM world, or run inside the WASM container at 2x or more the runtime/energy cost. And for what benefit?
It's one thing to get a game engine running in a web browser. Neat hack / potentially decent way to ship a client.
It's another thing to try to repackage existing working, relatively well engineered, server runtime systems inside it for almost no benefit at all.
TDLR: WASM is not a universal VM appropriate for server apps. It is a solution for shipping a certain kind of application in a certain circumstance. There are other, better, solutions for "containerizing" services.
Finally, after 25 years in this industry, the world I want to head to is higher level, where things are managed declaratively with explicit, visible, well described rules and relations and logic. WASM seems to me to push the other direction. Black boxes of fairly low-level code, each reinventing its own runtime wheel and with almost no visibility from the administration side of what's happening in there. I find that kind of sad.
I mostly agree with your overall point though. I’m not making the point that WASM right now (or even later) is the future of deploying backend services, however it is somewhat of a universal VM. If it’s a useful one is yet to be seen.
Putting it more clearly: Most services development is done in languages that have their own virtual machine (JS/TS, Go, JVM, Python, .NET). In what world does it make sense to run that VM inside another VM?
Finally, I think the experiences over the last 20 years around .NET and the JVM should have shown there is in fact not really such a thing as a truly universal abstract VM. A well-written VM tends to be written towards supporting the language(s) it is built for.
... Not unless you're willing to throw away almost all added value, and then you start looking like WASM. And then what's your value beyond native code, running on the hypervisor and/or in a container?
(There are in fact proposals for adding GC hooks in WASM. I'd have to spend some time reading up on them to evaluate whether they address my objections.)
b) Kubernetes is the platform as well as the application server.
- IDE code completion
- Can be machine generated/updated via the GUI management administration and graphical tooling on IDEs
Good luck doing that with YAML.
1 and 3 are actually foundational to how k8s works.
Forgive me for saying this, but I’m getting a “i don’t want to invest any effort understanding anything and so Kubernetes is bad” vibes from these comments.
It’s not a particularly pleasant experience to discuss anything with you, as after you make a particularly vapid comment that is naturally rebuffed you seem to just try to make snarky replies rather than engage.
Please understand that if you post your hot takes here they may be discussed and challenged, and if you don’t want this then I would refrain from initially commenting.
In response to your comment: They do. All Kubernetes resources are typed with JSON-schema definitions. Because of course they are, how else would kubernetes validate anything. https://kubernetesjsonschema.dev/
Anyone who’s used k8s at all knows this, if only from the error messages. From this you get autocompletion and a wide ecosystem of gui configuration tools that work with everything, including custom resource definitions, which is really cool.
I used to like lens (https://k8slens.dev/desktop.html), but now I use the k8s plugins from IntelliJ
Which I don't use hence why I wasn't aware of it.
> The Dunning–Kruger effect is a cognitive bias whereby people with low ability, expertise, or experience regarding a certain type of task or area of knowledge tend to overestimate their ability or knowledge
https://en.m.wikipedia.org/wiki/Dunning%E2%80%93Kruger_effec...
That’s legitimate but, especially for open source developers, overkill and it pulls in a ton of maintenance work (e.g. I have several apps with a ton of CVEs in components which were never used but can’t be removed without breaking things, and thousands of lines of XML configuration which has to be analyzed to look for intentional customization to upgrade to the release version which has been patched. Now, maybe that’s doing it wrong but that’s multiple separate Java specialist development shops so it’s not just me being dense.).
[1] https://github.com/vmware-labs/webassembly-language-runtimes...
I really with the WebAssembly community would stop quoting that tweet from Solomon Hykes. It's taken somewhat out of context and while at the time it was a big shoutout to the wasm underdog it's now a bit over used. Wasm really needs to justify itself with more than just a tweet and I think that it can.
"Polyglot - 40+ languages can be compiled to Wasm" is a bit misleading which anyone would see if you listed the 40+ languages. A lot of them are going to be very obscure ones and a lot of the popular ones are on that list. Factually correct but you're just setting people up to be disappointed. "40+ languages oh boy!.....What the heck is Zig? and no Python?!" (yes, I know you can sort of run Python if you run the entire interpreter)
In fact GCC used to have a JVM target, called GCJ. It was removed due to lack of maintenance.
The browser engine developers effectively paid a huge up front fixed cost and now anyone can use Webassembly in nodejs which then spills over to more and more use cases like cryptocurrencies using a modified WASM VM for their zk sync layer two solution.
The JVM simply wasn't built for these use cases.
A generic compilation target to the JVM would have worked, but it wasn't available 15 years ago, and GraalVM is mature about now, but so is WASM.
Which also involves rewriting literally everything to Java.
> In fact GCC used to have a JVM target, called GCJ
Which again, had a Java frontend. So, nothing like targeting WASM with C/C++/Rust.
JVM can only run apps written for it: Docker & WASM don't have that limitation.
> it can run non-WASM and non-JVM apps like PostgreSQL, etc...
but WASM can run Postgres
- Article: https://supabase.com/blog/postgres-wasm
- Repository: https://github.com/snaplet/postgres-wasm
- HN Thread: https://news.ycombinator.com/item?id=33067962
Of course it does; WASM is just a format, still need to create the right functions that the runtime will call to do the useful stiff
Of course they do. You can only run apps on WASM that have been compiled to WASM bytecode. You can only run apps on Docker that have been compiled to whatever bytecode is supported by the container runtime (which can be x86, ARM, x64, etc.).
> but WASM can run Postgres
WASM can run WASM-compiled Postgres. It has to be specifically compiled for WASM, which also means it generally needs to be ported to WASM first (as WASM runtimes have a lot of limitations that arbitrary C programs probably don't conform to).
Someone put in the effort to get Postgres to compile for WASM. That's great :) Maybe someday every application will compile to WASM as the preferred choice over the linux interface.
Compiling apps for different targets is VERY MUCH not a simple, low effort task though. Something like a database that must have an incredible number of optimizations in the way it makes syscalls, will have to a full stream of work to keep each target running well.
It can be done. But if "one of these thigns is not like the other" with your three things- Docker is the odd duck out.
To get an impression of the performance penalty, just run the following query:
SELECT SUM(i) FROM generate_series(0, 1000000, 1) tbl(i);
This simple query completes in 100ms locally on my laptop, but takes 17265ms in postgres-wasm. That is a slowdown of 170x.Now that is not WASM's fault - when running the same query in duckdb-wasm [1] on my laptop the query takes 10ms using WASM, and 5ms when run locally, with a slow-down of only a factor of 2. But in order to achieve those results we did have to adapt the DuckDB codebase to compile natively to WASM. That is absolutely possible but it does take engineering effort - particularly when it comes to larger older projects that are not designed from the ground up with this in mind.
Seems like some X can now run in wasm should come with disclaimer (includes Linux)
This is incorrect, there is a long list of limitations that your C/C++ code must conform to in order to compile to WASM. There's a whole section dedicated to this in the Emscripten docs: https://emscripten.org/docs/porting/index.html.
The chances your existing C/C++ app will compile to WASM and run correctly are much smaller than with Docker. However, the chances your WASM-compiled code will be able to run in a browser are much higher than with Docker (which is the real "killer use case" IMO).
> not like the jvm which is designed primarily for Java in the front end
This is also wrong, JVM bytecode is explicitly designed to be polylingual and is the compilation target for many non-Java languages like Scala, Kotlin, and Clojure. WASM being a compilation target is not what makes it unique from the JVM.
LLVM bitcode isn't all that portable though, and of course, the binary will still be OS specific because C/C++ code relies on native APIs. The primary reason to do this is so the Graal JIT compiler can optimize native code and higher level dynamic script/bytecode together and remove interop overhead.
However, you can theoretically run whole programs this way inside the Graal sandbox. If you do that you get an emulation of POSIX that is reimplemented on top of the Java standard library, so the code becomes portable, and in managed mode there's an additional party trick - the native C/C++ malloc is replaced with garbage collected allocations and memory accesses are bounds checked. So code run this way gets all the memory safety errors blocked automatically. This upgrade comes with two costs though, one is slower execution/more memory usage, and the other is you have to buy GraalVM EE. The community edition can run bitcode, but not in the sandboxed/managed mode.
Oh and GraalVM can also run WASM. So you can have cake and eat it, everything running together via their 'polyglot' interop system.
That has some limitations, but for the most part it works just like you would expect pthreads to.
https://emscripten.org/docs/porting/pthreads.html#special-co...
I suspect that still won't be as optimal as if you'd compiled the application for the target architecture in the first place, but I would suspect for most applications the performance will be relatively negligible.
For example, making sure build works accross different platforms and machines. There are too many ways that something may break, incorrect SDK versions, missing dependencies etc. Docker makes sure the OS (container) to be setup correctly to handle the build.
[0]: https://nixos.org/
Unfortunately Docker only works on Linux, since it's tied to specific syscalls. Those on other platforms (e.g. macOS) can only run it in a VM (e.g. the Docker Desktop application is built on top of a VM running Linux)
> here are too many ways that something may break, incorrect SDK versions, missing dependencies etc.
AFAIK Docker doesn't actually address that. It provides a "Dockerfile", which is essentially just a shell script; users still have to manage dependencies themselves, e.g. by having their Dockerfile invoke an actual package manager.
> Docker makes sure the OS (container) to be setup correctly to handle the build.
Containers aren't operating systems; they only need the desired executable, plus its run-time dependencies (e.g. libc).
And Windows. On Windows, Docker can actually create and manage Windows containers in addition to Linux ones.
macOS just doesn't have the namespacing primitives for such a scenario as far as I'm aware.
The reason why I’m bullish on WASM is the hope that there will not be a “Linux container” or a “x86 binary”, but only universal binaries and libraries. Unlike the JVM though, they’ll be sandboxed and don’t impose a GC with 200 tuning parameters. Lastly, there is language interop on an FFI level between any two languages that bind against WASM. In short, it’s an interop dream of mine.
That doesn’t mean I endorse premature shilling of everything WASM. That can do more harm than good.
For server scenarios, no, this just isn't it.
The JVM does a lot of very useful things that WA does not try to handle. Java is not going anywhere and is still going to need a GC. The WA runtime does _not_ handle memory safety inside of the app at all as of today...
Nope. Most of the Dockerfiles I've seen will do wildly-unreproducible things, like running apt/pip/yum/npm/mvn/sbt/etc. without even giving any version numbers (let alone expected hashes)
This seems to be pretty rampant; for example, here's some AWS documentation which encourages such reckless behaviour (even giving the '-y' option to apt-get!): https://docs.aws.amazon.com/AmazonECS/latest/developerguide/...
Citation needed.
Well, I know about BSD jails and Solaris (later Illumos, etc) zones. How easy were they to deploy to an average cloud? How easy was it to reproducibly build and distribute them?
Or what else would you offer as a better docker alternative from 10 years ago?
They weren't, like at all (yes, I have tried them). The dockerfile for repeatable (enough) image builds and the simple command line for running a container without having to mess with making a config for some init system is really the killer features of docker.
>Citation needed.
Well if you actually read the whole sentence instead of bluescreening in middle of reading then deciding to comment on half of sentence that changes it meaning
> the rest of industry just did better before he was able to capitalize on it.
you'd maybe figure out that I was talking about k8s and such picking a container format and ditching the rest of things Docker made. Not stuff that came before.
Docker as a company got relegated to "a repository" that they decided to monetize so people started going around that too.
The point of Docker is the ability to take the existing Rube-Goldberg-machine configurations of software, in any and many languages (including the gluing bash scripts), and put it basically unchanged into a controlled, isolated, replicated, shippable environment, with zero performance penalty.
It's very unlike WASM / WASI approach which requires recompiling stuff, runs non-native code, and completely changes the environment in which the code has to run. It's also like 2x as slow, compared to native code. It has its important upsides, but they are very unlike Docker's, in my eyes.
Secure means you can run arbitrary untrusted code, and webassembly cam do that, and docker can't.
This meme that containers are inherently insecure just because Docker doesn't attempt to be a security product needs to die. Docker hasn't been the only player in the container runtime space for a long time.
Thus, when people bring up "Docker is insecure" I try not to get down in the weeds arguing about the specifics, and instead point out alternative projects that are designed with security in mind. I find it's a much stronger counterargument.
Secure runtimes are a superior solution to virtualisation and separate kernels. There are only two secure runtimes in common use - for javascriot and webassembly.
The point about hardware support is very unfair - if you invest the same level of effort and hardware support into webassembly, you can also i prove its performance. Thats like compaining that electric cars suck because there are no chargers - its just infra.
Docker does not provide and security or isolation.
To have security and isolation with Docker you must use something external like SELinux or AppArmor.
Hope that helps.
Regards.
How is that not isolation?
Docker's value proposition is convenient. reproducible, self-contained packaging of software. It's the ability to deploy pieces of existing, battle-tested, gnarly and imperfect software next to each other, and care not about their conflicting or missing dependencies. It's more like Flatpak or AppImage, only more popular and easy.
This packaging also includes a kind of network insulation, exposing only the desired ports, making it easy to have VLANs between containers that do not interfere, etc. This is, again, not a serious security mechanism, but more of a convenience, but a very valuable convenience.
Congratulations, you have just discovered the additional value that WASM will bring to the Docker approach.
WASM is great, but it solves a different problem.
Docker doesn't solve any packaging problems, though. It just piggybacks off of other package management solutions and allows ad-hoc, unmanaged modifications to OS images (the convenience) and contains that result for easy distribution.
But nothing in that process ensures reproducibility— the package managers wrapped in Dockerfiles are typically non-deterministic: what any set of commands for them will do depends on the state of the internet at the time they run. Similarly, composing Docker layers is not like composing packages: reuse of packaged objects is minimal rather than maximal, granularity is course, dependency management details may vary from container to container (as they may be based on different Linux distros or language-specific package management ecosystems), and it's very easy to end up with software and configuration installed with which there is no associated package management metadata.
Docker doesn't know anything about packages. Container scanning tools that do things like produce a bill of materials or scan for known vulnerabilities inside Docker containers simply have to guess at what distro is installed inside the container and then reconstruct that information in a distro-specific way to the best of their ability! Docker doesn't solve package management issues so much as punt on them (which is, of course, convenient, because package management is hard).
> [Docker is] more like Flatpak [...] only more popular and easy.
I don't think this is a sound comparison, either. Flatpak is a desktop-oriented containerization solution, for packaging graphical software that will predictably need to interact with the local filesystem, GPU, sound, and other resources. It's also a solution that tackles security updates and deduplication in a serious (and effective) way involving some discipline and enriching shared runtimes with actual metadata rather than just composing filesystem layers together.
It may indeed be easier to crap out something which will be considered a valid Docker image than it is to crap out something which will be considered a valid Flatpak application, but that does not make Docker easier. It's not easier for a desktop user to keep a collection of 50 Docker containers patched for security fixes than it is for a desktop user to keep a collection of Flatpak applications patched for security fixes. It's not easier to take a random Docker container and plug it into your operating system's native file picker or sound system for use with graphical applications than it is to do so with a random Flatpak application. It's not easier to determine what the heck exactly is actually installed in a Docker container than it is to see what is in a Flatpak container. It is not easier to plug a random Docker application into your operating system's default password manager, and so on, and so on.
Writing Flatpak packages requires actually thinking about things that Docker doesn't because Flatpak actually solves package management issues (and other things) that Docker doesn't.
I would be curious to hear what is wrong in my comment above from anyone who has actually worked on general purpose packaging (e.g., written a package to be included in or overlaid onto a ports tree, maintained RPMs built from RPM spec files, run their own Ubuntu PPA, etc.), implemented tools that scan containers (e.g., SCA scanning or SBOM generation tools), or done reproducibility research.
Would someone with an awareness of full-fledged package management solutions based on or built with containers (e.g., Luet, Distri, Flatpak) really argue that having fine-grained abstractions for reasoning about dependencies or shared runtimes, performing security updates, etc., makes no difference as to what kind of software we're talking about and what problems it solves?
To me it seems obvious that
- not all software distribution mechanisms are package management solutions
- Docker cannot see or reckon with individual packages
- the Docker ecosystem relies on rather than replaces package managers, build systems, etc.
and so on. Are there serious arguments to be had here about those things, or do people just feel like my earlier comment was somehow unkind to Docker?Docker does handle deduplication on a certain level: every layer is only built once, and shared among all images that use it. This can be strategically used to seriously reduce the summary size of your containers.
Desktop users are not the target audience of Docker, except if you consider running a sham prod configuration on your dev machine desktop use. Containers are intended for the server side, and they are fine there.
Not sharing too much, and plainly embracing the existing chaotic practices of software creation and containing them, so that they don't interfere with each other, is the core value proposition of Docker containers. They do not require you to change your existing key practices at a lower level; your Babel / CMake / pyenv / whatnot setup can remain. But it changes the deployment story of it.
This sounds like a recipe for disaster to me and is why I haven't gotten into Docker.
If the software being deployed is too complicated to build and install without Docker, but Docker doesn't provide secure isolation, how can you be sure that this "gnarly and imperfect" mess of a system is secure?
What docker provides is a way to shrink-wrap a given build and all of its runtime dependencies.
Despite popular misconception, what docker does not give you is a deterministic way to build that software. A Dockerfile provides the RUN steps necessary to build the software, but dependencies must still be fetched over the network, introducing non-determinism.
You can spend 6 months happily using the latest release of some image, only to find that there's a critical bug or vulnerability that needs addressing ASAP. It kind of sucks for that complacency to turn to terror when you try to patch the software and rebuild, only to find inscrutable errors due to an absolutely bonkers build system.
"ERR: Version A of Foo is incompatible with version B of Bar"
Okay... but what happened here? What versions were we pulling before? Oh, the precise version isn't pinned in the build system, so.. I dunno. Great. It would really help if I knew if it was A that updated breaking B, or the reverse.
Then multiply that by 1000x.
And then add in (for Debian-esque base images) Apt repositories disappearing over time, git feature branches being deleted, tarballs falling off the edge of the internet, etc.
Now, before someone says "well, that's on you if your Dockerfile obscures so much build non-determinism!"
I agree with that statement! But that is a non sequitur with respect to the original premise: the build systems (and the web of dependencies they pull in) in third-party software you don't have ownership of is getting crazier and crazier, and Docker helps perpetuate this state of affairs, and the industry suffers as a whole.
I read an interesting thought: reproducibility is a spectrum. Docker isn't as reproducible as nix, but when used with version control and ci/cd, is damn more reproducible than zips with code and ftping them to servers.
You can't.
But it's not like escaping a container is going to happen because of a simple bug. You need an exploitable vulnerability in the containerized app that creates a path to escaping the container.
But yeah, if you want to isolate an app for security reasons, then you need a VM.
So what I'm trying to say is that making complicated applications easier to deploy doesn't seem like a win unless you also mitigate the increased security risk that comes with more complicated applications.
Sure, it doesn't reuse the kernel, but it's not a VM either.
This claim is dubious for many configurations of hardware accelerated virtualization. The hardware creates another ring 0 for each guest kernel, and guests run at the same level as the host. It's true that layering things like filesystems and networking incur overhead, but it's just as easy to pass through a physical disk, and bridge virtual TAP interfaces to physical NICs.
Hardware virtualization is very flexible, and there's a configuration out there that will meet the performance requirements of the vast majority of projects.
Consider VMs like those which run Javascript, or WASM. JIT compilation can get pretty close to C performance, as JVM and LuaJIT show though, given enough RAM at runtime, and money for development.
And for those use cases, where secure containment of running applications is not a goal at all (beyond maybe taking some basic precautions), I fail to see any value to recompile it to WASM, CLR, JVM, Z80 bytecode or whatever else.
Those who need isolation - yeah, WASM could be a very solid alternative to having a VM. But it's a pretty niche use case.
I just wonder how many founders and co-founder fail to understand the crucial value proposition of their business. I suspect one attribute of successful start ups is that over time they come to understand that aspect. And perhaps when we hear of startups "pivoting" that is not the result of a though process but instead a forehead slapping "why didn't I see that?" moment.
If you need to invest significant effort into dumping it, it's almost certainly cheaper to just pay for it. Especially so if the alternative makes any sort of compromise on developer experience.
If it takes two engineers half a year to get a replacement working you're already back in black, and honestly I'm not even sure why (in our specific case) it would take that long when there already are free alternatives.
And this is only the first example I saw. Now we have to root out all apps everywhere across the company that might integrate this tightly with docker.
Then there are performance considerations. Docker Inc apparently did quite a bit to improve performance, especially disk perf. We need to verify all of our existing workflows still work reliably.
None of this is hard, it just takes time and effort.
You had it for free for a long time? Lucky you!
The problem is really more one of "we operated so long without paying and now, blam, everyone pays next year".
It'd be sort of like if github deciding "You know what, everyone now needs to pay $7 a month/user for github". Perfectly within their right, but also a little bit of whiplash for a large number of people.
The next question is if this will last. There's already competitors to docker (rancher/podman/various k8s on my box things).
It seems that their assumptions were correct.
I do agree it is a product that is potentially worth paying for. I do NOT think that the product is worth the amount we were suddenly forced to pay. It felt like extortion.
But that last time was on Intel Macs; maybe the situation is completely different on ARM.
I mostly use the CLI so I have Podman a try and it took less than a single lunch to completely replace Docker for me. Performance is excellent for ARM, and acceptable for the x86 containers I infrequently use.
https://gist.github.com/acdha/9be1c3521af4f18d9f86264a889581...
Do you have links to recent benchmarks? My understanding is that there were investment in bridging the gap recently.
(Comparisons with small input sizes are not informative; look at larger runs.)
As far as I know, wasm won't let me apt-get install a bunch of stuff, set up cron jobs, glue together miscellaneous bash scripts, and magically run it without containers or VMs on any host architecture. That's the use case I'm most familiar with for Docker; wasm as I knew it was just a neat way to run untrusted native code on arbitrary machines with reasonable performance and security.
I think the disconnect is that Docker is also used as essentially a glorified package manager (sort of like Snap), combined with a runtime/interface that makes it convenient for devops purposes. From that perspective, I suppose there's not much difference between running a standalone binary inside its own dedicated Linux environment, and compiling it to wasm to run directly on the host OS, so long as the API remains unchanged.
It's just needlessly confusing to go as far as to call wasm a wholesale replacement for Docker. It's like saying Java is an alternative to Windows.
Well, Docker is not good at this, your Dockerimage can be as non-reproducible as it gets, it just pushes the problem to a different level. Nix and other package managers are the actual solution to this issue.
Are we intentionally not thinking about RAM usage in this dystopian world where we celebrate WASM-Docker progress without thinking of the drawbacks: memory inefficiencies?
However, it's true Wasm is not on that point yet. There are open threads about deallocate Wasm memory [1]. However, I expect these features, as well as Garbage Collection [2] will come to the stardard over time. This will allow modules and runtimes to properly manage memory usage.
And since one need an external process in any case, native containers wins as they are faster by factor of two over WASM.
EDIT:
It does not even make sense to use WASM inside a native container as an extra security layer. With the overhead of WASM one can just put a container inside a VM and still run things faster.
If your threat model includes hardware bugs, then a container doesn't really help, no? You can't really trust your containers without sandboxing them, and then you're killing your performance anyway.
And you're also comparing decades of VM and container investment to a handful of years of investment in WASM. WASM code today will run faster by a huge margin in a few years as the compilers improve.
But moreover, most folks don't care about hardware bugs letting untrusted code break out of a sandbox. Bugs have been letting code break out of VMs, even, for years. If a hardware bug is discovered, you install the microcode update or kernel patch and move on.
Which is to say, the performance isn't the reason for choosing WASM. It's good enough, in many cases. Being able to write a hundred or two lines of code to get pretty-fast and pretty-damn-secure sandboxing without needing to waste your time setting up and maintaining an elaborate breakfast machine of VMs and containers is the draw.
It promises to neutralize the playing field like Java promised, and Docker.
I’ve seen WASM do some cool shit, don’t count it out. Just factor in the irrational exuberance.
In addition with so many compile to JS technologies and chains, WASM is sort of another choice. Not a big deal for a team to choose it.
I don’t know enough about will it replace docker. But a lot of docker use cases are a bit of a leaky abstraction over what you are trying to achieve. For example why do I need to know what Alpine Linux is in order to run a node app? OK there is a node image that hides this detail but barely, you end up having to think about this sort of stuff.
Does it have the marketshare to make AWS make significantly different decisions? Remains to be seen. I guess people paid money for serverless. Could go that way
Edit: I never thought node.js would take over half of the information sphere, so read me more like a graybeard who is too young to be one.
In case you need to install anything in it as well as your application.
> Unfortunately, one of the challenges of running Docker on macOS or Windows is that these Linux primitives are unavailable. Docker Desktop goes to great lengths to emulate them without modifying the user experience of running containers. It runs a (light) Linux VM to host the Docker daemon with additional "glue" helpers to connect the Docker client, running on the host, to that VM.
https://mirage.io/blog/2022-04-06.vpnkit
WASM and WASI would actually be a cross compatibility story because you could compile to an architecture and syscall interface that is platform agnostic.
The issue is that I cannot easily compile.
Surely it replaces/is an alternative to images, not containers? If I have a wasm binary, there's still value in specifying the environment in which it runs, volumes it has access to, networking, etc.?
> Each traditional container gets its own control group as in docker/ee44.... On the other hand, Wasm containers are included as part of the podruntime/docker control group and one can indirectly observe their CPU or Memory consumption.
If I'm getting this right, WASI is basically just POSIX for WASM. This means that it does not provide some level of sandboxing that - for example - Deno has done. When running a Deno program, you have to actively allow network access or write access to the disk. It uses the built-in stuff from V8 for that.
Any idea why they did not include these kinds of permissions in the WASI standard? It seems like WASI was not designed to be run without some sandbox.
For me, the most interesting part is the component-model. It's still a proposal, but it will allow developers to specify the permissions for other modules (libraries) a main Wasm module may use. With this, you can give access to a folder to a module and that module may call another one without giving them those permissions. In other ecosystems, any library used by the "main" logic gets the same permissions.
If I have a complex program with a lot of dependencies.
What happens if one of the dependencies suddenly tries to write to `~/.local/dep_name/cache.raw` and I did not explicitly allow it to do so, since I didn't know it needs that location? In docker it would just create that folder in its own volume and the volume is deleted after the container is removed (if the volume is not named).
If I understood correctly from your comment, each WASI-runtime does not have a corresponding filesystem/volume.
But what does it do then? Will it simply crash?
How about Kubernetes, in other words, how does this scale (I understand the single process proposition, but can't see how it replaces multiple containers, which might be deployed on multiple VMs / hardware)? In other words, what is the WASM runtime running on? In the article they show a WASI layer, but that does not replace a VM / container (AFAIK), so you still need an OS to run on. I'm a bit puzzled.
EDIT: let me rephrase my question. In a docker container you can have your libraries and dependencies independent of the host system (eg. in your container you need libc 1.0, whereas your host system has libc 2.0). This is possible because the docker container does actually contain a copy of libc 1.0 if you set it up so. But in the case of WebAssembly this is no longer the case (this is what makes it possible to have smaller images).
But then you need not only the kernel from the host, but everything else around it. Unless your code does not depend or anything, or you compile ALL your dependencies to webassembly, which sounds interesting - I'm not saying it is not possible, but is this how it should work?
The "normal" shim on Linux is the runc shim (io.containerd.runc.v2). On Windows the shim is called runhcs (io.containerd.runhcs.v1).
The docker solution mentioned in the article modifies the "wasmtime" shim from https://github.com/containerd/runwasi so that it uses "wasmedge" instead. It also happens to be using an unreleased version of dockerd.
So how does this work with compose? Currently you need to specify the runtime for the container, there should be an option in the compose yaml for this.
How does it work with pods? You need to configure containerd's cri config with a runtime handler that specifies the wasmedge shim. Then you add a RuntimeClass to k8s and add that to your pod spec.
WebAssembly also allows for powerful constructs like the component model, where components written in even different languages can interact between them.
I often see a lot of technical articles posted to HN that are probably very interesting, but they assume the reader is living in the author's head and jump right into the details with little to no context.
-- February 2002 issue of MSDN Magazine
https://learn.microsoft.com/en-us/archive/msdn-magazine/2002...
What happened with .net is that C# is first class, F# is second class, and everything else is third class citizen at best (when not directly attacked via patent litigation).
https://www.microfocus.com/en-us/products/visual-cobol/overv...
https://www.silverfrost.com/1/default.aspx
https://www.eiffel.com/eiffelstudio/screenshots/
People do pay money to target it, go figure!
Library boundaries are not often so rigidly clear cut as to be a security boundary, ignoring also the performance & compatibility issues that come with such a thing.
Also good that it's open source right from the start.
Reproducible builds, consistent dev environments? I always thought it was to have production and development environments the same, but these statements contradict that..
Unless they expect you to run WASM on your servers..
They do, this is an emerging idea.
> I always thought it was to have production and development environments the same
This seems to be the first thing many people try to do with Docker. IMO it's actually not a great experience. In dev, you need to make changes, in prod, you shouldn't, so they're not the same at all.
Docker has many strengths: reproducible builds, consistent deployment strategies(k8s doesn't care what's in your container), a consistent DSL for building apps, The ability to extend a huge collection of other Dockerfiles to get what you want. I'm sure there are more.
In theory WASM looks really great, but the last time I looked at it, the gap was big enough to be a concern.
Similarly WASM will always lag behind the state of the art for native code (eg, new SIMD or other accelerated instructions). That's the price of portability after all.
That said, from what I've read, it looks like starting up a new WASM instance is pretty fast, so some places are using it for when they need to spawn up tons of instances all at once, without having to wait for a whole process to warm up.
I get the core idea of compiling other languages to WASM, but at the end of the day they have to talk to the outside world somehow right?
So it sounds like it's a new capability-based syscall interface, but they've ported big chunks of libc (specifically musl) to that interface so that a lot of things work.
I don't think I get this.
Does anyone have a pro/con list for docker and wasm at the server?
Is there a "Use Docker when..." or a "Use WASM when..." style guidance?