I’m curious how the impact affects development, deployment, etc.
I’m curious how the impact affects development, deployment, etc.
YJIT is pretty much transparent in production, if not it's likely a bug.
When we tried MJIT in production to compare it against YJIT, it causes lots of request timeouts on deploy, because the JIT warmup would take 10 to 20 minutes and it's much slower during that phase.
But YJIT warms ups extremely fast and with a much lower overhead, it's seemless on deploy.
The only thing you may need to tweak is `--yjit-exec-mem-size`, it defaults to `--yjit-exec-mem-size=256` (MB) which is not quite enough for larger apps.
As for development, it would work, but with code reloading enabled, you'd likely exhaust the executable memory allocation pretty fast, because for now YJIT doesn't GC generated code [0]. It will come soon, hopefully before the 3.1.0 release, but that's one of the reason why it's not enabled by default.
If you don't have the free RAM for it and you app is small enough, you can try lower values.
Also note that currently YJIT fully initialize that memory to make it easier to debug, so it's not even virtual. That's another thing that is on the roadmap to be improved soon, and ultimately you'll only pay for the part that YJIT really use even if 250MB is allocated.
But overall all JITs trade increased memory usage for faster execution, you have to store that generated code and metadata somewhere.
Other than that it's been rock solid and offers decent speedup across a wild range of benchmarks and real world applications.
So my personal expectation is for it to be enabled by default in 3.2, but I might be wrong.
In general it's not a big change. You should only turn on YJIT for long-running jobs -- the prod Rails server, plus probably background workers if you have them. YJIT does nothing unless you turn it on with the --yjit flag.
YJIT's going to affect memory usage, so it'll change your optimal number of processes and threads for your Rails server - play with it for your app specifically, because the percent speedup and mem expansion vary a lot from app to app. YJIT works fine with Puma and Webrick. I don't think anybody's tried it seriously with less-common servers like Falcon or Thin, but I'd expect it to work -- file a bug if it doesn't, because it should.
YJIT does speed up Rails - we're seeing about a 20%-25% speedup on little "hello, world" Rails apps (see: https://speed.yjit.org/), and about 11% on Discourse for a single thread. We don't have good multithreaded or multiprocess numbers yet, but it's in the works. YJIT scales with multiple threads/processes just like existing CRuby.
(All speed numbers are accurate as of right now, but may change over time.)
I would probably not use it in dev mode. While YJIT has pretty good warmup numbers, Rails throws away all existing application code for every request in dev mode. That's going to make YJIT a lot less useful. Play with it -- maybe I'm wrong for your app, especially if you use a lot of non-reloaded code (e.g. methods inside gems.) But I wouldn't expect great results in the development RAILS_ENV.
For deployment the short version is: add --yjit as a command line parameter in production mode and (probably) for your background workers. You can do this with "export RUBYOPT='--yjit'" or use your local preferred way.
Where I give weaselly-sounding qualifiers like "probably," that's because it can depend on your config. If you have plenty of memory but are often CPU-bound, YJIT is usually good. If you have really limited memory and your server CPUs are mostly idle, YJIT is usually bad. YJIT is also currently x86-only and runs on Mac and Linux but not Windows.
There has been a lot of historical demand for a Ruby config for servers: something that uses more memory to get faster operations, and that is optimised for long-running processes. YJIT is aimed directly at that. Non-JITted CRuby is mostly the opposite: fast startup, modest memory requirements, doesn't get significantly faster over time.
Have you tried it with Unicorn?
Of course since each of the unicorn process will generate its own executable code, the memory usage difference with Puma is even bigger, and copy on write can't help here.
Ideally we would be able to eventually do the forks of a "mature" unicorn child an hour or so after all is JITed...
JITed code can be invalidated and recompiled, so your forks would still drift over time.
I'm a big proponent of unicorn (and forking setups in general) for various operational reasons, but I think JIT might be the last nail in the coffin.
$ ruby -v
ruby 3.1.0dev (2021-11-08T09:35:22Z master 7cc4e147fc) [x86_64-darwin21]
$ ruby --yjit -v
ruby 3.1.0dev (2021-11-08T09:35:22Z master 7cc4e147fc) +YJIT [x86_64-darwin21]
$ ruby --yjit -e 'p RubyVM::YJIT.enabled?'
true
$ ruby -e 'p RubyVM::YJIT.enabled?'
false