WebVM: Server-less x86 virtual machines in the browser
leaningtech.com
leaningtech.com
I didn't see any benchmarks on the linked to page. I tried their sample Fibonacci program, but up to 100000 and ONLY timing actual execution (using the time Python module) to not include startup time, and WebVM only took 6.7 times as long as native for me. That's very impressive.
There's a similar open source project called https://copy.sh/v86/. Using their arch Linux image with the exact same Fibonacci benchmark, it take 44 times as long as native.
It's about ~10x faster on webvm.io compared to copy.sh/v86 and only ~20x slower than native, impressive stuff
I'd love to try v8 there, so we can benchmark the WebVM v8 against the Native JS in the browser... all using the same engine (is a bit meta, isn't it?)
Edit: I'm trying to run the following benchmarks [1]:
function mySlowFunction(baseNumber) {
console.time('mySlowFunction');
let result = 0;
for (var i = Math.pow(baseNumber, 7); i >= 0; i--) {
result += Math.atan(i) * Math.tan(i);
};
console.timeEnd('mySlowFunction');
}
mySlowFunction(8); // higher number => more iterations => slower
Results: 99ms in my Chromium browser (v8 JIT enabled), it breaks in the WebVM (after typing `node` and enter) with `TODO: FAULT af5147bf / CODE da d9 83`Results (not using console time): 99ms in my Chromium browser (v8 JIT enabled), 680ms in WebVM (1050ms in the first run).
Conclusion: WebVM node is only 7 times slower than native v8 in my local machine (macOS M1 Max)
https://blog.stackblitz.com/posts/introducing-webcontainers/
Try:
function mySlowFunction(baseNumber) {
const startTime = Date.now();
let result = 0;
for (var i = Math.pow(baseNumber, 7); i >= 0; i--) {
result += Math.atan(i) * Math.tan(i);
};
console.log(result, 'time:', Date.now() - startTime);
}
mySlowFunction(8);
I get ~100ms in Chromium and ~900ms in the WebVM (if I close the devtools)Python is also a bit tricky because it does things with pointers that I believe are hard to optimize (or maybe not, who knows!). Have you tried other languages/programs?
Does WebVM solve for workload transparency, CPU overutilization by one tab, or end-to-end code signing maybe with W3C ld-proofs and whichever future-proof signature algorithm with a URL?
DuckDB can query [and page] Parquet from GitHub, sql.js-httpvfs, sqltorrent, File System Access API (Chrome only so far; IDK about resource quotas and multi-GB datasets), serverless search with WASM workers
https://github.com/phiresky/sql.js-httpvfs :
> sql.js is a light wrapper around SQLite compiled with EMScripten for use in the browser (client-side).
> This [sql.js-httpvfs] repo is a fork of and wrapper around sql.js to provide a read-only HTTP-Range-request based virtual file system for SQLite. It allows hosting an SQLite database on a static file hoster and querying that database from the browser without fully downloading it.
> The virtual file system is an emscripten filesystem with some "smart" logic to accelerate fetching with virtual read heads that speed up when sequential data is fetched. It could also be useful to other applications, the code is in lazyFile.ts. It might also be useful to implement this lazy fetching as an SQLite VFS [*] since then SQLite could be compiled with e.g. WASI SDK without relying on all the emscripten OS emulation.
https://github.com/jupyterlab/jupyterlab-google-drive/issues...
So e.g. curl doesn't work without (File System Access API,) local storage && translation of e.g. at least normal curl syscalls to just HTTP/3?
The IP network is entirely virtual.
Now we could create a /dev/dom virtual device, and write dynamic web pages in pure bash. I love this.
echo "...." > /dom
Update the <title> tag: echo "TITLE" > /dom/html/head/title
Change the charset: echo "EBCDIC" > /dom/html/head/meta[1].charset // second <meta> tag
echo "EBCDIC" > /dom/html/head/1.charset // second child of <head>
Even go full XPath and replace a tag's inner HTML: echo "<div>abc</div>" > /dom/[@id='myID']
This is a horrible idea... // given:
<div id="myID">
// `id` attribute is located in
/dom/html/body/…/div[n].id
So unless you look at the contents of the files, you wouldn’t be able to find a certain ID. Because of that, DOMFS (when XPath is enabled) would expose that same "file" at `/dom/[@id='myID']` as well.I guess you could do something like this?
grep "myID" /dev/**/*.id
But why would you even use this “DOMFS”?I believe it will be possible to achieve similar state in the future just using Native Wasm/WASI (so no transpilation from x86 -> Wasm will be needed), but we are far from it given how slow the WASI standards move.
The shell is impressive: https://webvm.io/ (only downloads ~5Mb of resources for a full Debian distro)
By the way, it's spelled "Cheerp", with a lowercase p :-)
> Copyright (c) Leaning Technologies Limited. All rights reserved.
https://github.com/leaningtech/webvm/commit/6efab7e60bf6f173...
Assuming they wrote their own xterm interface (no idea if they did, I got as far as that), seems everything open-source is fetched by the client at runtime. This feels to me more like a bootloader than an OS. Not sure where that lands it license-wise whether merely linking to the image requires appropriate licensing and attribution but either way the work seems pretty straightforward to replicate assuming you have / can supply an xterminal-esque interface and can compile your OS image appropriately.
I don't think they're doing anything wrong licensing-wise but I guess it depends on how the law defines including software as a library, whether that needs to happen at compile time or run time, or whatever. Seems like a grey area?
Implementing the Linux ABI ourselves gives us the opportunity of a tighter integration with the Web platform anyway.
Any chance there could be a version with all the assets in one thing (say, GitHub Pages)?
Right now I see a ~6x slowdown on small-code benchmarks like sieve, but it goes up to 50x or more for large code like GCC. For QEMU it's roughly 1.6x and 2-3x respectively, so it seems like your JIT is slower than QEMU's.
In essence, for this approach, would the x86-on-x86 performance hit be similar or very different than the x86-on-ARM performance hit?
I wonder if this can be used to create a semi-decentralized website where visitors automatically served a vm to run, turning them into an edge server to offload requests from other visitors. The more active visitors, the more edge servers you have. Infinite scaling on the cheap! The visitors may not like you abusing their browser though, but there might be use case where this is acceptable, such as popular community run websites that too expensive to run due to huge amount of traffic.
You can buy ads. Ads cost about as much as ads, so you can buy and sell the same unit, and then run some compute for free.
I ported k(5) to html+js (ecmascript) during Iverson College (I think this was 2014?) and used webtorrent to connect to secondaries to run a scale test. The cpus are cheap, and they are slow, but it was a lot of (distributed) fun. I pitched the idea to KX (and a few others) to sell compute for fractions-of-pennies-per-hour but I think it was still a little early.
Do you think Now's the time?
The idea of running vm inside an ad to harvest compute from unsuspecting visitors... I think this might accelerate widespread use of adblockers even more if you successfully deploy this in the wild because people will notice ads are getting heavier.
Exactly.
> Is that actually allowed by ads provider?
On some ad networks it's prohibited by ToS and (in some cases) a review process, but it is extremely difficult to prevent in practice, especially if you have any understanding of how this works. I estimate perhaps as much as 50-90% of Google's adsense revenue comes from this, so they aren't (directly) incentivized to stop it.
> The idea of running vm inside an ad to harvest compute from unsuspecting visitors... I think this might accelerate widespread use of adblockers even more if you successfully deploy this in the wild because people will notice ads are getting heavier.
Perhaps, but people also have a lot of idle cores, so if you don't block networking and you monitor system performance carefully to ensure you don't affect things, for the most part people simply won't notice.
Quite a few people have been caught out doing wasteful things like trying to generate "coin", which definitely doesn't help, but there's also some interesting applications that have been run on volunteer-cpu-time (folding, seti, etc), so it seems plausible with some charitable examples and some care to avoid impact, this might be a doer?
To my knowledge the only way that one browser can talk to another is through peer to peer webrtc, but that requires a handshake.
I suspect that running arbitrary computations on peer browsers is uncommon today because verifying the output of an RPC executed on hostile machines is nontrivial unless you’re serving static files with a known hash, mining crypto, or solving an NP problem. I guess you’d also need a quota system to prevent users from taking over your botnet’s CPU time by spamming peer browsers with expensive RPCs.
Is there anywhere that documents the covered/uncovered syscall surface?
The more likely case is you get corporate and military clients who adopt this for security (ya know, load a known 0-state image) to check email (which is loading some old version of outlook from the image), and it ends up taking the entire workday before you can briefly use your system.
Basically giving a take on the recent https://www.airforcemag.com/fix-my-computer-cry-echos-on-soc...
But running docker container does not need to use dockerd. One can flatten the filesystem and have systemd run the container directly.
Note I don’t like to put my Chromebooks in developer mode or do the crouton stuff or whatever the latest is on that front..
EDIT, nm it has no tcp stack or outbound connectivity
https://stackoverflow.com/questions/14080845/tunnel-any-kind...
Would it be possible to compile GNU/Linux to WASM as a target platform? What's missing for that?
It's either a server somewhere handling tasks in a queue or the client is the server. I hate that we have to care so much about "words" but I've seen far too many walking away with an impression that the server is unimportant or requires little maintenance in this model. I'd argue it becomes even more important because of how persistence, consensus, availability, etc works /rant
That being said, always excited to see work coming out of webassembly. Network is the computer
If you'd like to access a server from another tab / iframe of the same browser, that's almost possible already, just some UX work would be required.
If you'd like users of the same page (or a separate specialized page) to connect, that could be possible with WebRTC.
If you'd like arbitrary hosts to connect, that would require a server side proxy, there is no client-side only solution that I can see,
So probably good for unprivileged non-networking applications.
id: /lib/i386-linux-gnu/libpthread.so.0: unsupported version 0 of Verdef record
id: error while loading shared libraries: /lib/i386-linux-gnu/libpthread.so.0: unsupported version 0 of Verneed record
bash: [: : integer expression expected
which would repeat no matter what command I try to run: user@:~$ ls -l
ls: /lib/i386-linux-gnu/libpthread.so.0: unsupported version 0 of Verdef record
ls: error while loading shared libraries: /lib/i386-linux-gnu/libpthread.so.0: unsupported version 0 of Verneed record
Very curious.Turned out my hitting F5 (because things appeared to have frozen - I assumed resource downloads had hung, I have flaky internet) corrupted the local ext2 filesystem in a way the runtime didn't detect and catch upon reload.
That's actually kind of interesting because HTTPS nowadays almost always means authenticated encryption, ie a guarantee that application data has absolutely explicitly not been modified. Thus, correct data was downloaded, but somehow corrupted on its way into IndexedDB, due to my hitting F5.
I'm very curious as to why that's happening, but I could only blindly speculate as to why.
Considering the bigger picture of the many block requests shown in the devtools, it would seem the ext2 block layer just transmits remote requests for small spans of "exactly what's needed right now", with no buffering or preemptive readahead. Given this runtime's nature as a likely-embedded component running specific bespoke applications (I get the impression this entire product exists solely or largely to run a retargeted version of the Linux Flash player) it is likely possible to optimize this quite effectively.
In fact, because of the A (then) B (then) ... (then) N serially blocking nature of the problem space, it could actually be quite straightforward to implement a fairly effective solution: if the block layer were modified to understand "here are the bytes you asked for... please also cache this random unrelated list of chunks to cache as well", the web server could then be modified to aggregate s=...&e=... requests, find correlations between ranges, and then preemptively send data estimated to be relevant for correlations over a certain threshold.
If you further modified the client to also implement a tagging system in the cache denoting "directly requested" or "preemptively sent by server (when requesting s=...&e=...)", when application code goes to request a given arbitrary block from disk, if that block turns out to have been preemptively cached and the preemption did indeed save a roundtrip, the runtime could send a pingback to the server to strengthen the relationship between the original request and the followup. (Given that this logic would be running on the server, and this approach provides a reward/feedback signal function, I'm almost tempted to wonder if a small neural network could be interesting to play with here, but it might be overkill.)
In any case, I'm 119ms away from the disk CDN... which might help explain my motivation to do the above. My ADSL2+ is hit and miss; today it's doing 6Mbps, give or take (especially for small high-latency transfers which take time to ramp up). The devtools is saying a full reload (clear IndexedDB, hard reload) takes 33.75s.
During that time it's entirely hung.
Some sort of early-boot progress-bar (an extention of the logic that puts the introductory text on screen?) would probably be a good idea. Maybe extend the runtime with a "userspace reached" opcode (or some sort of equivalent magic noop) called at the very bottom of /etc/bashrc, which hides the progress bar.
Lastly, a bit of an awkward problem that is Kind Of Interesting™ but doesn't have any straightforward next steps: while attempting to identify how long a cold start takes here I got into a bit of a fight trying to delete the IndexedDB with the site open. Naturally that didn't go down so well, with the app recovering and auto-recreating the DB. Somewhere in there between closing the tab and trying to reload everything, I realized the browser process had started using 100% CPU, and my laptop was at 94°C. Interestingly, everything was still entirely responsive (indeed the only reason I'm submitting this was because the form autosaver still worked; this was slightly too many butterflies at once to reason about and I completely forgot to save my post text), so I quit Chrome via the menu and... it didn't fully exit (I had to hit ^C twice (I run Chrome from a terminal) to make it fully quit). All the threads shut down but the root browser process was still yelling at my CPU. gdb decided it was stuck in futex_wait_cancelable(). Yay. Apparently the IndexedDB has a few handle leaks in it or something :/.
Errrr…. yes, “on its way”.