Sorbet: Stripe's Type Checker for Ruby
stripe.com
stripe.com
This article also made me laugh, because it reminded me of one of my small pet peeves about the Ruby codebase at Stripe: the fact that you would often find `merchant`, `account`, `invoice`, etc used as method parameters that represented the _ID_ of the resource rather than the resource itself. So Sorbet definitely helped with that, but it also could've been nice to just write `invoice_id` instead... :P
Makes me nostalgic though, good times!
Your critique of dynamic languages adding types and failing would be more understandable if you mentioned some examples or explained what Typescript got wrong.
I also think Stripe’s API (external) should not be moving ids and objects. Given some payload in which ‘account_id’ is always present and ‘account’ may be the object (using ‘expand’ IIRC?) or not makes a lot more sense to me.
1. Ambitious team who wants types does work to get the initial version passing in CI. Importantly, it's only checking at `# typed: false`, which basically only checks for missing constants and syntax errors.
2. That initial version sits silently in the codebase over a period of days or weeks. If new errors are introduced, it pings the enthusiastic Sorbet adoption team; they figure out whether it caught a real bug or whether the tooling could be improved. It does not ping the unsuspecting user yet.
3. Repeat until the pings are only high-signal pings
4. Turn Sorbet on in enforcing mode in CI. It's still only checking at `# typed: false` everywhere, but now individual teams can start to put `# typed: true` or higher in the files they care about.
5. Double check that at this point it's easy to configure whatever editor(s) your team uses to have Sorbet in the editor. Sorbet exposes an LSP server behind the `--lsp` flag, and publishes a VS Code extension for people who want a one-click solution.
6. Now the important part: show them how good Sorbet is, don't tell them. Fire up Sorbet on your codebase, delete something, and watch as the error list populates instantly. Jump to definition on a constant. Try autocompleting something.
In my experience trying to bring static types to Ruby users, seeing is really believing, and I've seen the same story play out in just about every case.
One final note: be supportive. Advertise one place for people to ask questions and get quick responses. Admit that you will likely be overworked for a bit until it takes off. But in the long run as it spreads, other teammates will start to help out with the evangelism as the benefits spread outward.
Luckily I think Stripe is a well-functioning org with smart and reasonable leaders—much easier to present that case than e.g., trying to win debates on static vs dynamic typing on the internet!
Nowadays, its extremely rare for me to come across projects that don't use Typescript already!
I've moved on from Ruby to Elixir and we use typespecs at my current job, but I still never do in any of my own code. Elixir does have a way to do type hinting which I do use and appreciate when appropriate. So I'm not "against" types and like the idea in theory, but I've never really felt the pain.
The most tangible argument I've bought is that types help out in a huge codebase. I can definitely see this. This is really solving a social problem, of course, because on smaller teams, it's much easier to enforce coding standards, especially if you follow XP and pair all the time. For example, in large dynamic codebases I've worked on, calling something `object` when it really means `object_id` would never fly. Variables and functions always have to be full words describing exactly what they do. Functions are only ever allowed to return one type. And of course everything is thoroughly tested. Of course I understand that not all these things are practiced everywhere and nor do they need to be. But for me that has always made for incredibly readable code.
I'm writing this just to give the perspective of a dynamic weenie and maybe looking to be convinced a little more? I also know three incredibly good programmers with 20-50 years experience each who all wrote in typed languages for years and then couldn't be happier not to have types when they moved to Ruby. So ya, I dunno what I'm expecting here, but mostly offering my perspective.
But the TL;DR: if I join your team and you want me to use types, I'll do it! But I really feel there are ways around it but those ways may not be for everyone (and sometimes perhaps impractical).
hmmm, I dunno, this was longer than I meant it to be, haha.
The thing I left out was that it's only been four month of being forced to write typespecs (I actually secretly resent it when I am forced to add them to private functions) so maybe I'll come around?
> I also know three incredibly good programmers with 20-50 years experience each who all wrote in typed languages for years and then couldn't be happier not to have types when they moved to Ruby.
Typed languages from 20 to 50 years ago were very inconvenient to use and had poor compiler diagnostics that made type checking a lot less intuitive, so this isn't surprising. Modern typed languages though are very different and should not be conflated with legacy or historical ones.
I use types at work in Typescript and Swift but honestly I'm not very convinced it's worth the hassle in most projects. I feel that if the build tool already can tell me what type I should be using, why do I need to pedantically add them everywhere?
@spec full_name(String.t, String.t) :: String.t
def full_name(first_name, last_name) do
"#{first_name} #{last_name}"
end
As someone who historically hasn't done any typing, I still find this a little noisy, but it lets me read the function as normal and can look at the spec only if there is any confusion.- type signatures offer succinct descriptions of what a function does - I write less unit tests with proper usage of types (if you need a natural number, take a natural number type, not an integer) - building your domain and giving names to structures that are in the application (passing around MailerOptions as opposed to an 'object that contains ...') - concrete interfaces for components/services/modules in my codebase
> it's much easier to enforce coding standards, especially if you follow XP and pair all the time.
IMO where it's possible to add process that makes the bad thing an error, I do that instead of requiring/trying to enforce discipline.
> or example, in large dynamic codebases I've worked on, calling something `object` when it really means `object_id` would never fly. Variables and functions always have to be full words describing exactly what they do. Functions are only ever allowed to return one type. And of course everything is thoroughly tested.
Agree on all fo this, but regarding tests, using types does prevent me from having to worry too much.
There was a point in time where writing lots of unit tests that tested non-user inputs (is the incoming thing a number? is it less than zero somehow) was seen as clever (and it still is a good idea where appropriate). These classes of bugs just don't exist in a properly typed codebase. If you need an Integer that can't be less than zero, you want a Natural (haskell parlance).
Excessive unit tests in untyped codebases are the result of not properly specified types. If you have a function that takes a certain type, there is no reason to check if you got a string (or not that type). If you pick a language with a top of the line type system, you don't even have to check for nil.
Realizing that nothing is actually safe in languages with nullable types was a big jarring point for me way back when. I think it was when Java introduced Optional<T> that I realized it should actually be everywhere, and all the Haskell I'd been doing on the side just made way more sense.
> I'm writing this just to give the perspective of a dynamic weenie and maybe looking to be convinced a little more? I also know three incredibly good programmers with 20-50 years experience each who all wrote in typed languages for years and then couldn't be happier not to have types when they moved to Ruby. So ya, I dunno what I'm expecting here, but mostly offering my perspective.
So I say the "convincing" part a bit lightly -- I know that people must come to their own conclusions, I can only suggest the idea. People in both typed and untyped languages can be incredibly productive (my friend is much like you, so productive and disciplined that types don't seem to add much), and I can admit that adding types requires more up front thought (good) and more ceremony (~bad, especially if you don't have type inference).
Maybe the big problem is that the really good implementation of types is in languages that can be hard to approach (Haskell, OCaml, other ML langauges). The second best IMO is Rust but that has the whole separate ramp of ownership/borrowing which is hard for different reasons. Go has some good concepts (protocols) but a bunch of other warts.
I find that writing software with a good type system gives me confidence that the code is correct. It's almost axiomatic because you're building the structures that are being checked/used by the code, but nevertheless I have that much more confidence in a Typescript codebase or a Haskell codebase than an equivalent untyped codebase.
Basically I think like this:
- If I want correct, I write Haskell
- If I want quick, I write Typescript
- If I want performance, correctness and efficiency, I write Rust
Also, even if you don't believe me at all, there is a trend -- python adding typing, sorbet and other systems for ruby, typescript gaining popularity, etc. The crowd sometimes has some wisdom.
I understand that types are helpful, especially within the context of learning new code (both as a new dev, and more easily understanding new code written by your team), but my hesitation to "jump ship" to them comes from weighing all the "boilerplate" code necessary to get them working versus the tangible benefits I feel they'd give.
To that end, maybe enumerating over the benefits for your friends (ideally more specifically than just "fewer bugs" or "easier to read/write") might tilt the scale more towards using them.
The article shows some good examples of problems that types solve, but many of them could also be "fixed" with better naming schemes, standardizations, and/or refactors -- and then wouldn't also be littered with type annotations (which take up 50% of the code in their example, [1], where I've also added a couple other easy-to-understand "solutions" with less cruft.
IMO, most of the typing syntax in their examples is distracting from the actual code -- especially when just writing better or more standardized code (which you still need to do with types) would make the typing syntax, again IMO, less necessary.
[1] https://sorbet.run/#%23%20typed%3A%20true%0Aextend%20T%3A%3A...
In Haskell for example, writing signatures is actually both unnecessary (they will be inferrable, most of the time so it's basically only for humans to read) and a net positive (they fit well prose wise and you can use them as documentation) but that's a harder hill to climb.
I do want to note that one big advantage of specifying signatures is that now your IDE can help you when using the functions and detect when something is going wrong.
Maybe "no more trying to figure out what the input/output of a function that you didn't write is supposed to be" would be a draw?
I've encountered a few Rails projects in the wild that do this. One solution is to make liberal use of the `to_param` method. This method converts objects to strings that are intended for use in URLs. Of particular note, it's the identity function for strings and numbers, but returns `.id.to_s` for ActiveRecord models. Using this within definitions makes your function polymorphic for whether it accepts a model or an id.
If you do this widely, would probably be best to monkey-patch in your own `to_id` method.
This obviously helps with things like `null`.
It completely changed the way I code. You have to think a little bit more about how you structure your code if you plan to hand it off to a theorem prover. I unlearned several bad habits (I unlearned even more with Rust).
If you have a type system then you should leverage it, not bolt on something extra
Have you considered using pre-commit?
But there's tooling (first-party and third-party) that will either download or generate RBI files defining constants that come from gems. `srb init` is the first party solution, and Shopify's `tapioca` gem is the most popular third-party solution[1].
Unfortunately, because Ruby doesn't have import statements at the top of every file, Sorbet can't just do something like silently treat unknown imports as not having a type (like TypeScript and Flow can do), because then it would never be able to tell between "exists but unknown" vs "typo; does not exist" for constant definitions. This definitely makes the adoption process a little tricker compared to other languages, but it's generally a one-time thing once you've got the tooling set up.
Also if you're ever having trouble getting the tooling to work, there's a lot of people chatting about Sorbet daily at https://sorbet.org/slack
Sorbet is something I've been interested in using for a couple years and finally got a round to actually trying it out. I tried to use Sorbet with ruby 3.1.1 but unfortunately it didn't "just work" which I think is crucial for mass adoption. I want to give the benefit of the doubt and say its my local env that causing issues with Sorbet but in a fresh `rails new test_app --api` project, I'd expect `srb init` to work without errors... maybe I need to give it another go, curious on your thoughts above tho! :)
I wrote up an FAQ about the state of Ruby 3 and RBS here:
https://sorbet.org/docs/faq#when-ruby-3-gets-types-what-will...
The tl;dr is that RBI files (not RBS files) will probably always be the preferred way to declare types for third party code (because it will always support exactly the same set of features that Sorbet does). We have some people in the community look into teaching Sorbet to read the RBS format, but the existing parsers for RBS files are written in Ruby and are very slow, and there are some ambiguities in the spec that make writing a third party parser that compiles to native code tricky. You can see an attempt to write a fast RBS parser in C++ here[1], but again given that RBI files do everything we need them to right now and we have other features people are asking us for, we haven't prioritized RBS support incredibly highly.
Sorbet works completely fine without RBS files!
Our general approach to metaprogramming at the moment has been two-fold:
- Use ahead-of-time code generation powered either by runtime reflection or ad-hoc static analysis to generate RBI files declaring things that have been metaprogrammed. - Build type system features, errors, and autocorrects that encourage people to structure their code in ways that doesn't require metaprogramming to solve.
Metaprogramming is definitely still a sticking point, but the existing solutions work ~okay and the rest of the upside Sorbet provides make it worthwile to power through.
Next challenges:
- Make it faster. While the post was talking about how fast it is, it wasn't telling the whole truth. Turns out some type checking operations in a 15 million line codebase are still slow, and we're working on making those faster.
- Add more IDE features. At the beginning of this year I put a lot of work into making Sorbet's parser more tolerant of syntax errors, which helps things like autocompletion work better. We also want to make more code actions, autocorrects, and refactoring tools, to bring Ruby in line with what you'd expect from other typed languages in the IDE experience
- Add more type system features. Shapes and tuples are a huge unimplemented feature still, and people ask about it all the time. There are a handful of other type system features (happy to list them if you're curious) that would also let people write idiomatic Ruby and still have good typing.
Lots left to do!
https://www.typescriptlang.org/docs/handbook/2/objects.html
(Ruby and JavaScript mean slightly different things by the word "object" so we chose a different word.)
Flow also has a distinction between exact and inexact object types:
https://flow.org/en/docs/types/objects/
where the difference is whether other, unspecified fields are allowed to hide in the object, or whether values of type `{foo: number}` must have only the `foo` field, and no other fields.
# typed: strict
extend T::Sig
sig {params(x: Integer, y: String).void}
def run(x:, y:)
puts(x + 1, y)
end
args = {
x: 1,
y: "Hi"
}
run args
type checks just fine in sorbet, but is an error in ruby 3+. It works though (with errors) in 2.7 and lower which is what Stripe uses from what I gather.You're right that Sorbet does not yet catch these bugs yet, and we'll likely get around to it in the future. It's being tracked in this issue[1].
We haven't quite prioritized this yet partly because it doesn't actually prevent you from using Sorbet in a Ruby 3 codebase, it just won't report all the errors it could (e.g., Sorbet allows using `T.untyped`, which can also cause runtime errors to go unreported).
Now that you've jogged my memory, there used to be some weirdness in our internal representation for method calls that made implementing this feature tricky. But we've since refactored some internal data structures and now it's probably a lot easier. Maybe I should look into how much work this would be again ...
Is that supposed to be how it works?
Couple of questions:
> The declare_method call above acted like a decorator on the def call method: it would check that the msg argument given to call was a String and that call returned a String on every invocation
How was this implemented? Dynamically altering Object#send or something along those lines?
How is the story for sorbet and vim/nvim?
Are there "run time" or "code gen" uses for sorbet? Like generating swagger/openapi documentation/schemas based on typed Api methods? Or vice-versa - scaffolding sorbet-typed Api from a swagger.json? Or something similar for graphql (or, well, SOAP..)?
The mechanism in the post is the same mechanism in use today with Sorbet signatures.
To get the `sig` method in scope, you have to put `extend T::Sig` in that class (or one of its parents). When `sig` is called for the first time in a class, it monkey patches that class to install some overrides of the method_added method. Ruby calls this method_added method for you every time it creates a method (from any means, static or dynamic). Code is here:
https://github.com/sorbet/sorbet/blob/master/gems/sorbet-run...
> How is the story for sorbet and vim/nvim?
Great, honestly better than VS Code for everything except autocompletion. I personally use Sorbet with Neovim.
> Are there "run time" or "code gen" uses for sorbet? Like generating swagger/openapi documentation/schemas based on typed Api methods? Or vice-versa - scaffolding sorbet-typed Api from a swagger.json? Or something similar for graphql (or, well, SOAP..)?
Not that I know of unfortunately, but I also pay more attention to the type checker and it’s bugs than the tooling people build around it to get real world work done
Also, it's fast! I'm in total agreement with the point made in the article. That makes a huge difference in developer UX.
I started thinking about this a bit and I came up with the conclusion that the single biggest difference between structs and shapes is really iterating over keys. I spent some time trying to create structs by which you could iterate over all the keys and all the solutions seemed clunky or inelegant.
Solargraph combines inference and insight from YARD docs (standard for many gems, plus Castwide has written more YARD for the standard library) to make some pretty good guesses. Crucially it has plugins that add the insights from popular gems with static analysis (e.g. reek, rubocop). I maintain solargraph-rails, which parses your Ruby to make guesses about (surprise) Rails.
The typeprof gem can help IDE plugins make typing guesses based on your tests. This project is interesting to me because it's going into Ruby 3.1 so I think it reflects awareness from the core ruby team that many programmers are not ready to add types to their code.
solargraph: https://github.com/castwide/solargraph solargraph-rails: https://github.com/iftheshoefritz/solargraph-rails typeprof: https://www.youtube.com/watch?v=UTMj51j9yEg
The alpha releases are also a big concern. We are stuck on a 300 commit (release) old build and can never upgrade safely.
We have also never been able to get the VSCode extensions to run.
Thanks for Sorbet, but I’d suggest people outside Stripe to look elsewhere.
Still, if it helps it helps.
[1] https://crystal-lang.org/reference/1.3/tutorials/basics/60_m...
Edit: Perhaps I spoke too soon https://blog.appsignal.com/2021/01/27/rbs-the-new-ruby-3-typ...
At least the JS crowd had the decency to buy into a whole new (far better) language instead of a bolt-on solution.
Though Java still has some great strengths, especially the 8+ functional programming features and the concurrency library is great. If I could use Rails with Java it might be a different story though, since I hate Spring.
What I'm realising now is that Java has (had?) a really clunky type system and that there are other languages that do types better, so I shouldn't use Java as my reason for avoiding them.
In my experience, large codebases of those types of languages have a lot of "magic" thing happen. There's a lot of implicit stuff that one has to guess or spend time "following the code" to understand what it is doing.
And I say this after having built a major lending platform from scratch in Ruby, including a major Machine Learning scoring system in Python, having to maintain with a good sized payment system in pure JavaScript, and nowadays dealing with a major trading/liquidity system in Ruby.
They are fun languages, but once the code and systems start to scale, static typing really helps. For that reason I've seen a lot of these endeavours try to move to TypeScript or other typed languages.
That's because the theory of gradual type systems was only worked out in the '00s. Before that, you could have a static or dynamic type system, not anything in between. Common Lisp did have type annotations, but they were hints for optimization, without any guarantees. They were also local to subroutines only. Dylan[3] is an example of an early implementation of the idea, but Dylan was several years late and, without being able to compete with Java, died without ever being widely used.
The proper theory was first established by J. Siek[1] and W. Taha in 2006. It's distinct from nominal static typing which uses a single top type (like Object in Java) or generics, and obviously it's different from both purely static and dynamic typing. It took almost a decade for the idea to start gaining practical implementations - I think the original was a made for Scheme, and one of the first implementations was Typed Scheme for PLT Scheme, which continues on as Typed Racket[2] today. Typed Racket is unique in that it enforces the types even on the untyped side, by wrapping values and exports in contracts.
The idea proved to be useful in practice, and started being adopted in various (non-Scheme) dynamically typed languages, starting with TypeScript for JS and Hack for PHP. On the other hand, some statically typed languages also became gradually typed, most notably C#. The implementations continued to improve, shrinking the parts of their respective languages that could not be statically typed. In dynamic languages there are still features that cannot be practically expressed in static type systems - most metaprogramming and code generation falls into this category - but they are generally "good enough" for day to day coding.
Gradual typing is useful in the same way static type systems are useful: it can prevent certain kinds of errors by marking known-invalid expressions without the need to run the code (so, for example, can help you find errors even in code that's not covered by tests); it helps in writing tooling for the language (eg. go to definition, find references); it helps make the code clearer for the reader (no need to break into a debugger to see what kind of value a given identifier refers to); in some implementations it may also help in optimizing the runtime performance, but that's rare. The "gradual" aspect makes it easier to adopt when the codebase grows larger - the bigger the codebase, the more useful static types are, but by the time the codebase grows large enough to justify static typing it's too big to rewrite in a different, statically typed language.
In short: writing small projects or prototypes in a dynamically typed language is faster while maintenance and expansion of large projects is easier in statically typed one. Gradual typing lets you go from one to the other without a huge cost of a full rewrite.
[1] https://wphomes.soic.indiana.edu/jsiek/what-is-gradual-typin...
Actually, both metaprogramming and code generation blur the phase distinction between compile and run time - but this makes them a great fit for dependent typing systems, which do pretty much the same thing. So there's no reason why more refined static systems could not express these idioms.
On dependent typing and gradual type systems - I'm not sure they are possible to mix. I think none of the current implementations do anything remotely similar to dependent typing. It would be a really cool if it worked though. I'd love to be able to track length of a list in mypy type system (for example).
For the curious:
Right, because they lift values into type system, those values need to be well-typed themselves. You wouldn't be able to represent gradually typed dynamic/Any value then, I think. I might be wrong though, my understanding of type theory is limited at best :)
> to get proper optimization and reduce the overhead of dynamic checking and conversion
That's probably why very few gradual type systems are taken advantage of when optimizing. Typed Racked does this, but Typed Racket disallows untyped values within the typed module. Hack probably does it, too, but I have no experience using it. Common Lisp does it with correct optimization settings. Raku does this, but Raku was written from the ground up with gradual typing in mind. There are projects that try to transpile fully-typed subsets of dynamic languages to C (can't find the link right now), but that's a bit different. Other popular gradual type systems are basically advanced linters. There's a lot of development and potential in this area though, we'll see how it pans out in a few years :)
see for SBCL:
Types are useful when they don't get in the way. As are method contracts.
For key parts of the code, there was no type safety where we expected it.
In the end, it felt far from being like Typescript, we opted for removing it, instead we added some runtime type checking and we document with YARD. Far from ideal, but that's the tooling available.
The gem integration is terrible currently: we wrote the gem, fully typed with sorbet, but for some reason the type checking was completely ignored in the main project where we referred it
The TLDR for me is: I’d still be willing to keep using sorbet, if issue number 4 below (that the LSP isn’t very responsive) would be fixed. Otherwise, it adds more work than it removes from my workflow, so I've stopped trying it out.
Start with the negatives: 1. It’s a lot of grunt work to set it up properly in an application like ours with many dependencies. Specifically, sorbet and/or related tooling tries to generate RBI (equivalent of typescript’s index.d.ts) files by actually importing and running your code, and doing introspection on the types of the arguments of functions. I managed to work around this problem, but it’s a reflection of how young Ruby typing is as a whole that this step was a time-sink for me.
2. There’s no mechanism to do the equivalent of yarn add -D @types/react-table (which would install typings for the react-table package). You basically have to copy paste from this github repo[0] manually.
3. Some really popular gems still don’t have types. For example, IIRC, devise’s typings are either non-existent or are uselessly incomplete.
4. The sorbet LSP isn’t very responsive, at least in neovim. I asked @jez about this just now, so hopefully I'll get a response.
5. Super verbose syntax.
Positives: 1. Thanks to this repo[1], there’s actually a way to easily generate typings that would cover a lot of the dynamism of rails, it works quite well.
2. It’s legit helped me catch errors with my code.