Ruby 3.2 preview 1 with support for WASM compilation
ruby-lang.org
ruby-lang.org
The syntax feels a little rough although I have no ideas how to make it better:
Regexp.timeout = 1.0
...
/^a*b?a*$/ =~ "a" * 50000 + "x"
I think I would favor the: long_time_re = Regexp.new("^a*b?a*$", timeout: 1.0)
version instead but I use the `=~` almost entirely, so that would still be a big style change. Probably end up setting a global timeout per app and then overriding for individual checks as needed?Adding a timeout is a bit strange, first because you don't know in advance how long it's going to take for large search. The timeout is a failsafe against something that should be fixed in the first place.
the point is the language is broken fundamentally and because its syntax sugar there isn't anything you can do realistically. a function call can be replaced, deprecated, and then removed and all of that can be automated and be non-disruptive over a few releases.
Their suggestion is essentially making stack overflow a feature in regex, & then allowing that stack depth to be tuned
best solution is to have an implementation like re2 that does not have those problems.
By design RE2 isn't fully compatible with Onigmo. As another poster mentioned, a hybrid "use RE2 when possible; fall back to Onigmo otherwise" approach was considered and rejected for well-explained reasons https://bugs.ruby-lang.org/issues/18653Maybe in addition to `Regexp.timeout = 1.0` there could also be a `Regexp.parser = :re2` option with `:onigmo` being the default.
What are your credentials for being you?
The OP is not _confused_, they just don't like the solution.
So yes, `Regexp.timeout` is supposed to be a default setting for the app, and when really needed you can override it with the `timeout:` key.
No easy way to override it locally when using =~, but I can't imagine too many cases where you would want to use a local timeout anyway.... can just switch away from =~ syntax for those.
This is mostly a denial-of-service mitigation tool, something you'd just want to apply globally to avoid disasters spawned by malformed or malicious input. In practice, it's hard to imagine a use case where you'd really want to be twiddling the knobs on a regexp-by-regexp basis.
I was initially thinking that it would make sense
to always ask yourself "how long should this take"
and tune appropriately, but for the vast majority
of regexes that's overkill
More than being overkill, it's actually impossible right?The execution time will also vary greatly based on base CPU performance, and current server load.
A regexp that takes 10ms to process right now might take 500ms tomorrow when your server is under heavy load. So we can't predict how much time each regexp "needs."
But, like you said, we can set a somewhat ridiculously high limit to help prevent regex-based oopsies or re-based DoS attacks from dragging us down =)
Would WAS compilation help solve the ruby / rails timeout problem?
https://devcenter.heroku.com/articles/h12-request-timeout-in... > H12 - Request Timeout in Ruby (MRI) > Rack-timeout limitations > Due to the design of the Ruby programming language’s Thread#raise API, which timing out code depends on, there is a known bug with rack-timeout that can put your app into a bad state if rack-timeout fires too frequently.
(Other related URLs) https://github.com/ankane/the-ultimate-guide-to-ruby-timeout... https://headius.blogspot.com/2008/02/rubys-threadraise-threa... https://www.schneems.com/2017/02/21/the-oldest-bug-in-ruby-w...
Edit: Ah I guess it’s just the WASM vm if it includes everything
In order to gather the dependencies into the tarball in the first place, they need to be installed. And not all your dependencies are handled by `bundler`. In any interesting project you're going to be depending on system libraries that Rubygems has no way to signal to the OS that it needs. You need to `apt install <foo>` for interesting values of `foo`.
That means you've got a few choices: either you document what those OS packages are and rely on the operator knowing that they need to pay attention, or you try to automagically install them when you unpack your tarball, or you include everything you need from the build system in the artefact. This means running the application and finding all the files it touches via `dlopen` and hoping you exercised all the interesting code paths.
Option 1: you're back to installing dependencies. Option 2: you're wishing you'd just used `dpkg` to build yourself a .deb, and you're still installing dependencies. Option 3: now you need to figure out how to swap out OS libraries after the ruby interpreter has started, and need to get all the configuration right to point into your untarred filesystem, and you're wishing you'd just used Docker.
The portability sounds great, but I’d like to see how performance is impacted. Edit: also curious about file size.
That said, I am not seeing a link in here about how to actually use this code. Is there a good tutorial/example somewhere?
[0] app.datastation.multiprocess.io
[1] github.com/multiprocessio/datastation
<html>
<script src="https://cdn.jsdelivr.net/npm/ruby-head-wasm-wasi@latest/dist/browser.umd.js"></script>
<script>
const { DefaultRubyVM } = window["ruby-wasm-wasi"];
const main = async () => {
// Fetch and instntiate WebAssembly binary
const response = await fetch(
"https://cdn.jsdelivr.net/npm/ruby-head-wasm-wasi@latest/dist/ruby.wasm"
);
const buffer = await response.arrayBuffer();
const module = await WebAssembly.compile(buffer);
const { vm } = await DefaultRubyVM(module);
vm.printVersion();
vm.eval(`
require "js"
luckiness = ["Lucky", "Unlucky"].sample
JS::eval("document.body.innerText = '#{luckiness}'")
`);
};
main();
</script>
<body></body>
</html>
[0] https://github.com/ruby/ruby.wasm/tree/main/packages/npm-pac...They also created an awesome playground to try Ruby online [1]... all powered by Wasmer/WASI [2]!
There were attempts in the past to use ruby in the frontend, by compiling it to JS (opalrb), so I looking forward to see how the ruby community exploits WASM.
While I don't know how good it will be it still interesting the possibilities that WASM bring to the table, no more forced to use one language for the browser, use whatever language you like!
Maybe it’s not the most popular project out there, but I don’t see a reason to flag it as a mere attempt.
No installation required :)
<internal:/usr/local/lib/ruby/3.2.0/rubygems/core_ext/kernel_require.rb>:85:in `require': cannot load such file -- socket (LoadError)
or is it because you're using docker buildx which emulates aarch64 on x86_64 through qemu?
Also, what's the impact on a typical rails/ruby dev? Do they have to learn anything new to enjoy improvements WASM brings, or will all the changes be 'under the hood' (i.e. in ruby and/or rails)?
WASM on the server is interesting for companies providing compute services like edge computing or traditional FAAS like AWS Lambda. As WASM has very good security features (sandboxed) it is much cheaper to provide a secure edge compute service for WASM only than one that supports Java, .NET, Node, Python, Go, ... However, neither the enduser nor the developer are the winner here. For most languages WASM is slower. For developers it is obviously easier to provide just your .NET code to the cloud provider than to compile it to WASM. Most (if not all) cloud providers also support NodeJS, so if you would need to develop an edge function I'd highly suggest just writing it in Typescript.
So, WASM can help cloud providers make even more money or for very demanding web applications. Otherwise, a lot of hype ...
https://www.destroyallsoftware.com/talks/the-birth-and-death...
Bernhardt in the talk explicitly mentions asm.js which is the precursor to wasm (it's even mentioned in the wikipedia article you skimmed a bit too quickly). asm.js was released Feburary 2013.
I'm surprised HN has such a short memory, but the impetus for that talk was a clearly disturbing trend at the time implying that everything should be done in javascript. Node.js was gaining rapid popularity, people were discussing javascript as the new C for using as the language to write example code in, and while things like asm.js were exciting, they seemed to point towards the hilariously nightmarish future Bernhardt is discussing there.
WASM requires an interpreter which must be native.
The argument is that this interpreter can be smarter about what crosses OS security rings. But those same improvements could be done in the native compiler or interpreter.
The next argument could be that many things using the WASM target would focus more effort on improving it so all WASM targets benefit outpacing their individual optimizations.
This one is harder to dismiss outright, but instead of optimizing for machine code you are now optimizing your WASM output.
Also this intermediate byte code representation already exists for both LLVM and JVM, which many languages target.
It is difficult to see WASM magically improving performance at all and especially not dramatically enough to encourage people to switch to it for that reasoning.
WASM isn’t faster than native code. It’s that an operating system written from the ground up to use a language VM (for instance, wasm) to implement all memory protection, on the system — and most importantly, no other memory protection, including page tables or processor-level isolation that normally separates kernel code from user space — may end up being more performant than what we have now.
Running wasm on a standard OS kernel like Linux/windows/darwin is not going to give you the benefits. The benefits come from eliminating system call overhead associated with switching in and out of the kernel protection context, which is something you need if you’re executing raw machine code which can load/store any memory address. If you simply eliminate the ability to run arbitrary machine code, you can just let everything run in kernel space and use the language VM (like wasm) to protect memory. The result may or may not be faster overall, but it’s purely hypothetical today because essentially no operating system works this way. (Microsoft wrote a research OS back in the ‘00s to try this out but it ultimately didn’t turn into a shipping product. There may be other research OS’s out there that toy with this idea, but nothing in production… maybe the old symbolics lisp machines work on this principal but I’m not sure. There may have been some similar machines in the smalltalk days as well.)
They called it SIPs (software isolated processes). You can see the benefit of WASM for this problem, though. It has gained significant traction, and it is significantly simpler than CLR. I really hope something comes of it.
> For instance, the SPIN Web server executes entirely in the kernel address space.
http://www-spin.cs.washington.edu/
But like you said, not in production.
the fact that i am more frequently having to dig up this tweet to show people lately is astonishing
(Somebody correct me if I'm wrong; I know what WASM is but I'm not sure how it's employed in practice outside of in-browser tech demos of games and things)
https://blog.cloudflare.com/webassembly-on-cloudflare-worker...
I'm not sure how the former would work in a pure Von-Neumann-architecture, where it's not possible to reference code as data (and consequently also not possible to emit new code).
What would be interesting to JIT compile in the context of Ruby isn't the Ruby interpreter itself (both AOT and JIT should be fine for that), but the Ruby programs it would be running.
Also Rails is plenty quick these days, tons of people running it at massive scale.
In the browser you have to also factor in the warmup time.
I'd imagine an interpreter will suffer a lot because certain C tricks like computed goto don't work directly. (This will hopefully be improved by future Wasm proposals)
(Note: that's still plenty fast enough for most use cases, and performance will improve)
It's more accurate to say that the average case is 20-50% slower. The best case is on par, or slightly faster than native code[1].
[1] Measurements from our original paper, https://dl.acm.org/doi/10.1145/3062341.3062363
Engines are even faster now.
A good analogy is that Opal is like PureScript, whereas Ruby 3.2 is like GHCJS.