Sorbet Compiler: An experimental, ahead-of-time compiler for Ruby
sorbet.org
sorbet.org
How does one generally get to work on compilers? Is it mostly through academia? Through open source contributions to compilers?
I've been meaning to write a blog post about this because I get asked about my personal journey to working on the team a bunch.
The makeup of the current team is a split of:
- people who worked on programming languages & compilers in academia
- people who worked on developer productivity/product infrastructure at large tech companies
- people who just really like types (me)
Which is to say: there's a deal of luck involved, but there are also some fairly straightforward ways to inch closer to PLs/compilers job in industry.
That all being said, the piece of advice I always give people who bring this up: there's nothing stopping you from working on PLs/compilers right now! Find a ticket labeled "good first bug," show up in the #internals channel on the Sorbet Slack[1], and someone would love to help you get started.
Don't care about Sorbet? That's fine! Pick your programming language of choice, find where the developers hang out, and repeat the process. Oftentimes, there's nothing holding you back from working on PLs and compilers except yourself.
You can just do it! As a place to start I'd recommend looking for tutorials on recursive-descent parsing, and just try to make a program that parses a toy language and transpiles it into C, or some other language you're familiar with. Then you can grow it from there.
You can get really fancy with things automatic parser generation, and using an LLVM back-end to enable optimization, but writing a simple compiler is a reasonable project for someone with even a basic handle on programming and a little creativity to just jump in and figure out.
If you start looking for mini languages in a normal project (or where they would fit) you will see that there are frequently many opportunities to do something useful with small domain specific languages.
I suggest starting with parser combinators rather than page 1 in the Dragon Book. They are a nice way to get things done quickly.
FB isn’t as steeped in modern service-oriented ideology as most similar companies, so that `www` monorepo represents a lot more code than you probably expect, including stuff you’d normally think of as “services”.
True backend services do exist outside the PHP application, and are typically written in C++. Java, Rust, and Python also exist in various niches.
The only mainstream language I have never seen anywhere at FB is Go, although I’m sure any language you can imagine is used by _some_ random team out there.
I’m not sure why Hack would be unpopular. On the contrary, it’s the default choice for starting new projects since all of the internal tooling is built mostly with Hack in mind, so it’s the path of least resistance for developers.
I don't think I'll use the compiler (for now, at least), but I'm very interested in seeing how it grows over time :)
For example, having to wrap values in `T.must`, `T.nil`, etc. is a net loss. I read Ruby code full of Sorbet annotations and I wish they were gone.
You're adding useless boilerplate to satisfy the tool, and you have to change your code in unintuitive ways so the tool can understand it.
IMO the Sorbet additions do impact readability, but it dramatically improves understandability. In a codebase past a certain size/number of contributors, you're always going to have ambiguity about what sort of data you're working with. For example, in my part of the codebase we have models representing an entity that is identified by a number in external systems and represented as a whole data object in our code. If I have `foo.entity`, it's unclear just from the "real" code whether that'll return a data object or just the numerical identifier. Having the Sorbet annotations on that method (and available on hover in the IDE) make it a lot easier to navigate the codebase and resolve these ambiguities.
On a smaller code base I may not bother with it for the reasons the original commenter mentioned.
The numbers and testimonials in that presentation are two years old at this point but I can only think that means they’ve improved since then.
I would love to read more if anybody can find it.
I think I understand GP's motivation: RBI files and RBS files are two different formats, and as a user of the language, people tend to want to use the officially blessed solution the language provides.
In case you weren't aware, parlour[1] is a popular open source project for working with RBI files. I believe it supports transparently converting between RBI files (Sorbet) and RBS files (Ruby 3).
There is also rbs_parser[2], a C++ parser for RBS files to convert them to RBI files, written by Shopify, a major user of Sorbet.
Stepping back: I haven't personally read many complaints from Sorbet users describing how the current state of RBI/RBS interop gets in the way of what they can actually do with Sorbet. Almost all the feature requests we get about Sorbet (both inside Stripe and outside) are for fixing bugs or implementing new language-level features. RBI files as implemented seem to work.
Sorbet already has an extensive set of RBI files covering the Ruby standard library (at least as good or better to my knowledge than any existing repository of types for RBS files), and there are plentiful tools for working with RBI files, listed here.[3]
If lack of first-party RBS support in Sorbet is holding you back from trying Sorbet, I'd strongly encourage you to give Sorbet a try anyways! Many people have shared great experiences adopting Sorbet in their Ruby codebases.
[1] https://github.com/AaronC81/parlour
- A type checker
- LSP-based editor tooling
- A compiler
But we haven't:
- implemented our own GraphQL, protobuf, gRPC, or JSON libraries
- built a Ruby debugger or debug protocol adapter
- built custom performance monitoring tools
- etc.
There might come a time when it makes sense to fork the language, but we've been very reluctant to do so from the start, because we know what we'd be throwing away. "Compatible with Ruby" has been an explicit design principle of Sorbet from the start:
https://github.com/sorbet/sorbet/#sorbet-user-facing-design-...
I agree it's too bad that Sorbet's syntax choices weren't adopted "officially" by Ruby Core – I think matz had some unfortunate strong opinions about types never appearing in normal .rb files – but I wouldn't be too disheartened a Rubyist as long as the community adopts Sorbet anyway, which it seems to be doing.
(disclaimer: former Stripe, former Rubyist)
Really no to that, where is this coming from? Stripe is Stripe and yes perhaps a couple of huge companies here and there may give Sorbet a try (although I believe Shopify didn't go this route, neither has Github afaik), but as a community Ruby is a dynamic community. And also, Ruby is much more characterized as a tech for small-mid teams (where types are arguably not helpful) than as an enterprise tech which is mostly typed.
I don't get the whole choose Ruby and sparkle types on it with some tool, you have much better tools to do that - they are called Java/C#.
And even the "better tooling" part is somewhat subjective. The added strictness and explicitness is valued by some, but for others it's seen as degraded readability.
It seems like he works on a different part of Shopify, which might explain our difference in opinion. The Ruby code I write has to be typed: true at minimum, usually typed: strict.
>Architected this way, the Sorbet Compiler turns Ruby into a language for writing Ruby native extensions! Instead of having to write C, C++, Rust, or some other compiled language to write native extensions, people can continue to write Ruby but gain the benefits of native compiled speeds.
To me this is the biggest feature that could impact the whole Ruby Ecosystem. Along with another Ruby JIT that is currently being tested at Shopify.
> instead of having to ship an entire language runtime to production
Except they ship the CRuby language runtime in production. Ruby is not a language that can run without a runtime.
> Not only did we not need Java VM-level interoperability
So they almost discard the entire idea because they don't need a specific additional feature?
> choosing either alternative Ruby implementation would have made for a difficult migration path.
So what is it?
> Stripe relies heavily on gems with native extensions
Yes that's a problem on JRuby (when there is no java extension in that gem), but TruffleRuby supports native extensions, as very clearly stated in many places.
> as you can imagine, a multi-million line Ruby codebase over time starts to depend on Ruby-the-implementation, not just Ruby-the-language.
Except all serious Ruby implementations know they need to be compatible with whatever CRuby does, not just an incomplete ISO specification of the language. In fact alternative Ruby implementations match CRuby behavior as much as possible for compatibility, even when it seems weird or makes little sense (they report it in this case but have to match behavior anyway).
TruffleRuby might not be 100% compatible with CRuby yet, but I would say it is [pretty close](https://eregon.me/blog/2020/06/27/ruby-spec-compatibility-re...).
> To adopt JRuby or TruffleRuby in Stripe’s most important services, we’d have to be able to run all the code or none of the code [of a service].
Well yes, to get significant performance gains, one needs to try new things. Don't they have a staging environment where they can do experiments?
Although this may have been a very old impression. Pandemic and lock down has an surprising effect on the sense of time.
- The Sorbet types are hints for optimizations in the Compiler. The compiler doesn't blindly trust them, but rather it checks whether they're correct and if so does something faster.
- The Sorbet compiler can frequently check types much faster than the interpreter could because it can look directly at the object representation, rather than having to fall back to calling a full-blown method like .is_a? or .nil?. Many common type checks are a single assembly instruction, so type checks are actually fast most of the time.
- Sorbet is and has always been designed to have runtime type checks.[1] These have been a part of Sorbet since even before we open sourced the typechecker. Every method already does runtime signature checking, when interpreted, and this is no different when compiled.
- The power of LLVM means that a lot of these type tests end up coalescing. For example, if a signature says "this method accepts ints" and then the compiler sees "if this is an int, I can do something faster," that's frequently only one type test, because the power of LLVM magically coalesces the checks.
It sounds like you're taking an approach fairly similar to what Facebook is doing with Hack these days: Dynamic checks implied by types help catch bugs and can be often optimized away, and static types are hints but not trusted, but the type checker means that using the hint is almost always a good idea. Is that accurate?
When I'm looking at a production performance profile result, trying to figure out why the compiler didn't speed something up, the first thing I do is add or improve types. It has never slowed down the resulting compiled code, and usually speeds it up substantially.
On the other hand, this won't be very attracting for small projects, where the biggest bottlenecks are in 3rd party frameworks and libraries. Not being 100% compatible with upstream Ruby (from what I heard), this compiler might not be able to save your day.
I'm curios, what's the baseline like here for Ruby? I.e. how does vanilla ruby compare to something like Go or Rust for a use-case like webservices?
Go is about 2x slower than C.
Rust is about the speed of C.
CPython is just like Ruby MRI. It does not do type checking or use type annotations and it's basically an interpreter not a compiler.
Also mypyc is still very alpha, while cython is production-usable for years now.
Sorbet is great - chuck it on your core entity's and core methods and boom it feels great. Thank you team for creating this tool in OSS and continuing to work on it. Looking forward to the future of it! Thank you Jez!
I understand if ruby (MRI) don’t want to make typing or ahead of time compilation mandatory —- I think that’s wise.
But would be neat if this would make its way into official Ruby eventually, so that projects could choose typing + ahead of time compilation if they wanted.
Just as Rails has influenced a lot of Ruby, I hope this will do the same!
Also interesting to see Mypyc, Pyre related AOT compiler (forget the name), and now this.
Great to see! We’d gladly compile our code in CI if it meant better perf (already do for the UI stuff)
Instead, I/O problems are frequently a language-agnostic set of problems, like:
- Oh jeez, I'm making a database roundtrip 2x more frequently than I'd like to! Maybe I should change my application to batch the database requests.
- Oh jeez, I'm running on cheaper EC2 instances without SSDs, maybe I should upgrade!
etc.
Granted, there's been a lot of interesting work done on Ruby concurrency lately, but it's far from a solved problem at the language level.
(background: former Shopify EM)
I'm currently working on Polyphony [0], a Ruby gem for writing highly-concurrent Ruby apps. It uses Ruby fibers under the hood, and does I/O using io_uring (on Linux, there is also a libev-based backend).
> It uses Ruby fibers under the hood,
I'm sure you are aware of ioquatix's work on Async/Falcon etc, how do you see your project differing from his work? And why hasn't anything changed in the Ruby server space - it's all processes and threads afaik.
I get it, this is a big one. How big though? I wish as a community we had a list of important gems that have an extension, maybe with a combined effort we can port them to the JVM or whatever else so people can switch between JRuby/Truffle/MRuby with relative ease. I know about the big ones: Of course Postgres/MySql gems come with extensions and probably most Ruby web servers. But what else - what are the big ones?
MRuby uses MGems and is just a different ecosystem entirely. It has parallel libraries but they're not shared with the above implementations.
After my experiments I came to the conclusion that type information would probably be better if specified in structured comments, something like TomDoc. That way type information would be had just by writing the documentation one should be writing anyway. (Two birds, one stone.) It would also make the type signatures very readable.
My beef with Sorbet is that it seems like these two goals have been conflated a bit. While it's true that they aren't entirely orthogonal to each other, you definitely optimize for different things depending on which goal is more important.
If I'm primarily interested in developer understanding, then it's really important to use a DSL for annotations that's easy to visually scan and parse. I don't think Sorbet is great in that department. Even among Sorbet advocates, I often hear complaints about the syntax. That's because there's a direct correlation between the expressiveness of a language (Ruby being very expressive), and the expressiveness of the DSL you need to describe that language.
Sorbet went with an approach that tries to capture as much of Ruby's expressiveness as possible. But the reality of day-to-day Ruby programs is often that you could capture 90% of the use cases with 10% of the DSL footprint. If your primary heuristic is developer ease-of-understanding, then leaving that remaining 10% language coverage on the floor could be your best option.
Something like Typescript has its cake, and eats it too, by virtue of becoming its own language -- it's able to be very expressive while still remaining relatively easy to scan and parse. On the other hand, Sorbet is fighting with one hand tied behind its back in having to stick to standard Ruby syntax.
A few years ago at Shopify, I worked on an internal comment-based type checker for Ruby that would instrument our methods at runtime, and compare their inputs/outputs to what was declared above in YARDOC. It was definitely skewed more towards developer understanding than program-correctness. You couldn't really use it for static analysis, but it was pretty good at providing guidance unobtrusively. It was just comments, after all -- you could color them however you liked in your editor.
Right before we started adopting Sorbet in our monolith, we had high-level discussions about which type checker we wanted to use, and we ultimately went with Sorbet. I think that was probably the right call (more cumulative momentum). But I do still wish that Sorbet optimized a bit more for ease-of-understanding by restricting itself to a simpler DSL (at the expense of some descriptiveness).
It's possible to write a YARDOC-to-rbi converter so that you can still rely on comments instead of having to pepper your code with Sorbet annotations proper, but I haven't seen much use of that in the wild. We used such a converter at Shopify for migrating off of YARDOCs in places, though.
But before I left, the strategy was to begin making inroads into typedness by attacking the most important places first (which in our case was often the inter-component abstraction boundaries). Teams were encouraged to add Sorbet annotations to their intra-component code where appropriate but it was not a blanket requirement.
Regarding JIT/CRuby/Sorbet compiler -- I don't think it's either/or. A company of Shopify's size can afford to fund work on multiple projects with overlapping goals.