Pitchfork: Rack HTTP server for shared-nothing architecture
github.com
github.com
As I understand it, Pitchfork is an implementation of a reforking feature on top of Unicorn, a single-threaded multi-process Rack HTTP server.
The general idea of 'reforking' is that in addition to pre-forking worker processes at initialization time, you re-fork them on an ongoing basis, in order to optimize copy-on-write efficiency across processes, especially for memory allocations that might happen after initialization. The Pitchfork author mentions sharing YJIT machine code as a particularly important use-case [2], one that has become relevant in recent Ruby releases.
I never saw much interest in the concept when I originally introduced the reforking feature to Puma, and it remained a kind of interesting experiment that never saw much production uptake. I'm thrilled that Shopify is iterating on the concept through this project and I hope the technique will continue to be refined and developed, and see broader adoption as a result.
[1] https://github.com/puma/puma/pull/2099 ; https://github.com/puma/puma/blob/master/5.0-Upgrade.md#fork...
[2] https://github.com/Shopify/pitchfork/pull/1#issuecomment-125...
As for why it wasn’t used much as a Puma feature I don’t know, but there are a few challenges with it that prevented me from putting it in production. The most important one being that new workers end up being grand children of the original process, so if the middle process die, you may end up with zombies etc.
It’s also very scary to fork a process that may have live threads currently processing a request. I believe Pitchfork solves most of that.
That said: Static files never should be chunked.
Chunking is ONLY interesting with dynamic responses.
Rounding out the previous thought, the idea with many forking servers is to boot up to the point before a request is served and then fork off for each request. You do gain CoW benefits, but if you have any lazy data structures that are reified in the call, now each child is faulting. Pitchfork will take a child that has processed a request and promote it as the parent, replacing the original process. Now, this new parent is the new CoW base with the expectation that forks from that will result in even greater memory sharing. For a framework like Rails, there's a lot that happens after a request is received.
A response might be a couple MB, a medium sized Rails app will need a couple hundred MB to load it’s code etc.
As for the concern about forking a process that may have live threads currently processing a request, this should already be solved in the Puma implementation. The worker shuts down and finishes serving all pending requests before reforking. There is also an `on_refork` hook for to trigger extra garbage-collection to maximize copy-on-write efficiency, or close any connections to remote servers (database, Redis, ...) that were opened while the server was running.
I wonder what that financial calculus looks like between these options for a large organization:
1. trying to basically reinvent the workings of an entire programming language and migrate your Ruby apps to these new tools/runtimes/typecheckers/whatever
2. incrementally rewriting critical paths to something more performant (Go/Rust) & just throwing more servers at the Ruby stuff you haven't managed to replace yet.
Tobias (Founder/CEO) was an early Rails core contributor. So Ruby on Rails is in Shopify DNA.
[1]: Koppel, Solar-Lezama - Incremental parametric syntax for multi-language transformation (2017) - https://dl.acm.org/doi/10.1145/3135932.3135940
Except without actually solving the problem of making the translation nice and readable. But this kind of thing at least gives you options for changing language.
Big issues off the top of my head.
Legacy support. Legacy code is already difficult to understand. If it was originally written in another language that would only make it harder to grok.
Transpiled output is often ugly. Developers in general seem to dislike modifying generated code for a number of reasons. I imagine transpiled code would end up implicitly frozen, and new code would only be added to new files.
Language idioms are hard to translate. Some structures are unique to languages and their equivalents may be considered bad in other languages.
You would either need equivalent libraries in new languages or also translate dependencies. You could probably find equivalent libraries, but their scope may vary.
Large scale code organization varies between languages. Things like Dependency Injection may not translate at all. Some languages may use config and some use code for the same thing.
I am assuming Ruby, used here meant Ruby Rails. Or specifically, ripped Rails out of everything as they scaled?
The calculus, in many situation would indeed be in flavour of another framework or language in ~2013. The cost of CPU Core has dropped significantly since then, we are not far from having 128 x86 CPU core in a single socket. Even if could only do 10 request per second per core at a target latency, this is 1280 Request per second on a single server. ( In the context of a non API serving App, or Server rendering App )
Edit: Miscalculation here, it should be 128 Core, 256 vCPU, so 2560 RPS.... but the main point still stand.
But at the Scale of Shopify, ~$33B Market Cap ( ouch.....I remember they passed $100B market cap in 2020 and $200B market cap in 2021... ) They could finally afford to pour some resources into Ruby. The only Top 10 programming languages without large cooperate resource backing. I am pretty sure the performance potential of YJIT and Ruby tooling has a ROI of less than 2 years at their Scale.
Facebook is probably the best example of this with PHP and Hack [1]
1. https://en.m.wikipedia.org/wiki/Hack_(programming_language)
This is not saying you shouldn't, it doesn't really matter to me. But I think a lot of the time this is driven by engineers who feel like they need to do this rather than a financial decision to save money
1. Company started as Rails monolith
2. Some components of the Rails monolith stretch Ruby capabilities beyond where "throw more servers at it" is still reasonable or effective
3. Factor out some backend functionality into dedicated services which can be scaled separately, Rails still serves as the API gateway calling these services. Anything without scaling issues stays in the monolith for now.
4. Eventually Rails is just an API gateway, no one in the org knows Ruby/Rails and its dependency management madness any more, and it gets replaced with a more-performant, purpose-built API gateway, usually something off-the-shelf.
Yes, I'm aware of 3x3 but for the most common usage of Ruby which if for Rails apps - there hasn't been anything like the gains PHP saw from 5.x to 7.x.
Especially from Microsoft, who has deep compiler/language expertise, and who owns GitHub (large Rails app).
Also, some performance improvements like JIT just don't matter that much for Rails because the primary problem isn't single thread performance it's the lack of cheap parallelism.
I think we're entering a period of increased experimentation and rapid evolution as demonstrated by projects like YJIT[1][2], improved inline caching[3][4] and Object Shapes[5] (also used by V8), and variable-width allocation[6][7], and smaller improvements like better constant invalidation[8]. Significant investments in TruffleRuby[9] are still going on by Oracle, Shopify, and other companies.
And recently, Takashi Kokubun gave a talk at Ruby Kaigi about the future of JIT compilers in Ruby that gives a peek at a whole new set of optimizations Ruby can work on (as well as some performance comparisons against other interpreted languages)[10]. You may be surprised to see how well Ruby (with the JIT enabled) performs compared to Python 3.
All of which is to say, I think there's quite a bit of performance improvement being made in recent Rubies, and that trend will likely continue for quite some time.
update And I forgot to mention that some very notable computer science researchers and their teams are working in the Ruby community now![11]
[1]: https://news.ycombinator.com/item?id=28938446) [2]: https://speed.yjit.org/ [3]: https://bugs.ruby-lang.org/issues/18943 [4]: https://bugs.ruby-lang.org/issues/18875 [5]: https://bugs.ruby-lang.org/issues/18776 [6]: https://bugs.ruby-lang.org/issues/18045 [7]: https://bugs.ruby-lang.org/issues/18634 [8]: https://bugs.ruby-lang.org/issues/18589 [9]: https://eregon.me/blog/2022/01/06/benchmarking-cruby-mjit-yj... [10]: https://speakerdeck.com/k0kubun/rubykaigi-2022 [11]: https://shopify.engineering/shopify-ruby-at-scale-research-i...
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
Please don't take my comments as unappreciation for the hard work going into Ruby.
Seems like we need benchmarks for actual web workloads. :)
There’s nearly an entire order of magnitude difference in performance between the best of PHP vs best of Ruby (~400k vs ~50k rps)
https://www.techempower.com/benchmarks/#section=data-r21&l=z...
https://github.com/socketry/falcon is an interesting project, but again, it's not clear how difficult it would be deploying a Rails app on top of this. My experience with these concurrency models in Rails apps is that one single gem could make a blocking IO call (for example) and negate all of the concurrency/performance gains. It would be cool if Ruby could be configured to throw errors for these types of calls to make finding and fixing offending gems easier.
There's a lot of really great projects happening and plenty to be hopeful about, but when that stuff will land or the changes the rest of the community and ecosystem should think about making still isn't clear.
That’s not really a use case we have, and the community is already investing a lot in that direction (ioquatix with fiber scheduler and ko1 with N:M threads and Ractors)
Every place I've worked in the last ~5 years has treated Ruby as "legacy" code for better or worse.
This has simply not been my experience. The jump in performance in the 3.n has been awesome, to the point were I even downgraded / removed servers.
Ironically the greatest speed (and memory) improvements have been in the swapping out the supporting JS ecosystem components that Rails uses, for example swapping Webpack for the Go based Esbuild.
I thought it was missed opportunities those working inside GitHub could not persuade Microsoft to at least budget out additional resources for Ruby. Although they might have their own battle trying to keep everything Rails and not moving to .NET
It does worry me a bit that all resources are coming from Shopify, hopefully they can keep growing so resources isn't a problem in the near future.
Where did I said Shopify not putting resources when the last part of my comment was exactly about Shopify putting resources into Ruby Rails.
Microsoft have contribution into Ruby? Or Do you meant Github has Core Contributor in Rails? Name me a single Microsoft employees actively helping in Ruby Core.
As a suggestion to the authors: A section titled "Why not Unicorn?" explaining the rationale behind this and why one would want to choose it over Unicorn, would be very helpful. If not in the current experimental stage, then in the "stable" one.
GVL contention is an issue when using threads of fibers, so it isn’t an issue with Unicorn
If you could choose a 50% reduction in process memory footprint, versus a 50% reduction in CPU cycles to serve your average request, which would you pick?
Memory is second. In 15 years CPU has never been a real concern.
I've often seen problems because of N+1 querying in particular, but maybe that's just the majority of codebases that I've seen.
> Memory is second. In 15 years CPU has never been a real concern.
In most stacks that I've worked in, memory seems to be the main cause of problems: be it a Ruby application, a PHP one, Python one or Java.
Though perhaps that's because larger servers are expensive and development environments in particular tend to be on the more conservative side of things. Either that, or the apps that I've worked with are good examples of Wirth's law and refuse to even run with less than 2 GB of RAM (though maybe that's more along the lines of commentary about enterprise Java apps in particular). When your development/test server has something like 8 GB of RAM or your computer has 16 GB of RAM and you need to run 7 or 8 of those apps (as well as IDE instances), you run into problems.
CPU is generally not an issue for most applications, though various databases and data stores instead. Perhaps it's batch migrations or just lots of in-database processing, or honestly just most cases where you use ElasticSearch - not only does it seem to eat as much RAM as you'll give it, but the same seems to apply to the CPU (especially when used as the data store for something like Skywalking APM).
On the bright side: most of the modern tech stacks are generally reasonable and there are often micro frameworks (or just more lightweight ones) to be found, which can be suitable for some use cases. Don't want Rails, use Sinatra. Don't want Laravel, use Lumen. Don't want Django, use Flask. Don't want Spring Boot, use Quarkus. That said, Node with Express.js and .NET with ASP.NET are pretty lightweight out of the box, which is nice. Of course, if you go down this route, you might end up with something that is better suited to a bare API, instead of fancy features like server side rendering/templating etc.
Most of our work happens in the background. We're running a Sidekiq process per CPU core with a concurrency of 20 — which, while the default, seemed to play the best during our tests. We set a max RSS of 2000 and use some code to automatically scale Heroku based on queue size. We do a lot of network IO.
It was a headache to get here, and a hole in our wallet, but I'm mostly happy with the application. Previously, we were running Resque in production on a super over-provisioned AWS node that was costing us $xx,xxx a month: Resque forks the app process for concurrency (we weigh about 450 MB); Sidekiq is a threaded model. Heroku, at least in this case, has been exponentially cheaper. We keep our RDS on AWS for at least some savings versus having it on Heroku. I cannot imagine what a xx-TB database would cost us at Heroku.
> what's the main benefit?
Drastically reduced memory usage by improving Copy-on-Write performance. For mid-sized apps it can even use less memory than puma depending on how you set it up. See this synthetic benchmark.
> Why did you decide to build reforking into a fork of Unicorn instead of contributing it to Puma?
Mostly for simplicity. You can build this in Puma, but there are a few extra challenges. For instance you don't want to refork when another thread is currently processing a request, you'd risk leaving some global resource in a corrupted state. So you'd first need to stop accepting traffic and wait for all requests to complete. The problem is, Puma doesn't have a request timeout, so if the request never complete, what do you do? Lots of small challenges like that. But I think it would be awesome if Puma was to take back that idea. Puma's fork_worker feature was a big part of the inspiration after all, and I'd even be happy to help Nate or someone else do it.
By comparison, Unicorn (the project upon which Pitchfork is based) is a single-thread, multi-process Rack HTTP server that does not buffer requests, so it's only designed for fast clients (or it needs to be paired with a proxy like Nginx to handle slow clients).
The narrower design of Unicorn results in a simpler, less-flexible potentially more efficient architecture.
The reforking feature that Pitchfork introduces to Unicorn was originally implemented in Puma [1], though they're controlled differently and the underlying implementations are entirely different.
[1] https://github.com/puma/puma/blob/master/5.0-Upgrade.md#fork...
this was the concurrency model for pre 2.0 apache. (i think in 2.0 you could pick)
often times you'd put a proxy in front, serving static assets with a lighter weight http so that slow loads wouldn't tax limited server capacity for big httpd processes that included full perl interpreters and all application code loaded.
so it seems it's almost exactly the same?
the models that replaced it were single process multiple threads and asynchronous (never block) models that didn't create or require one of many application processes per request.
i think the choice of the preforking model was a stability and portability thing. if one of the children crashed it was no big deal, where a crash in a single process application or webserver will cause the everything to go down. also back in those days threading semantics varied widely between operating systems, the apis themselves differed and linuxthreads were just really too new to run in production.