Ruff: A fast Python linter, written in Rust
beta.ruff.rs
beta.ruff.rs
I'm around if anyone has questions.
Personally I'm interested to hear from you what are the specific reasons you can't achieve this kind of performance with CPython.
Usually the major factors are: A) python's generalised data structures (int, list, etc); B) extra overhead of common operations like reading a variable value, calling a function or iterating over a collection; C) no real multithreading (i.e. the GIL); D) lack of control over memory management.
I'd love to know if there's anything else that makes Python that much slower.
1. The "fearless concurrency" that you get with Rust is a big one though. Ruff has a really simple parallelism model right now (each file is a separate task), but even that goes a long way. I always found Python's multi-processing to be really challenging -- hard to get right, but also, the performance characteristics are often confusing and unintuitive to me.
2. Ruff performs very few allocations, and Rust gives us the level of control to make that possible. (I'd like to perform even fewer...) We tokenize each file once, run some checks over that stream, parse it into an AST, run some checks over that AST, and with a few exceptions, the only allocations outside of that process are for the Violation structs themselves.
3. Related to the above (and this would be possible with CPython too), by shipping an integrated tool, we can consolidate a lot of work that would otherwise be duplicated in a more traditional setup. If you're using a bunch of disparate tools, and they all need a tokenized representation, or they all need the AST, then they're all going to repeat that work. With Ruff, we tokenize and parse once, and share that representation across the linter.
4. Again possible with CPython, but in Ruff, we take a lot of care to only do the "necessary" work on a given invocation. So if you have the isort rules enabled, we'll do the work necessary to sort your imports; but if you don't, we skip that step entirely. It sounds obvious, but we try to extend this "all the way down": so if you have a subset of the pycodestyle rules enabled, we'll avoid running any of the expensive regexes that would be required to power the ignored rules.
Biggest missing piece of the workflow now is a modern replacement for Tox/Nox; only one I see is Hatch, but that replaces a lot more, for better or worse.
If interested: CPython moved to a new parser, a PEG parser, in Python 3.9, and part of the motivation was to support language features like pattern matching, which introduced ambiguities in the grammar that the existing parser couldn't handle. For the same reason, they've been non-trivial to implement in the RustPython parser -- they either require clever techniques, or the parser needs to be written as a PEG parser. I am hoping to do the former, but I need to find time to prioritize it.
Love the tool
pre-commit run --all-files
because it can be much tricker or more error-prone to only lint the correct diff when you're dealing with MRs. I definitely notice the speed hit there, though it usually pales in comparison to running pytest.Honestly mypy is the slowest one now, which ruff doesn't intend to replace.
- isort (import statement sorting)
- pyupgrade (syntax upgrade for newer Python versions)
- pylint (general linting)
- pycln (remove unused imports)
- pydocstyle (docstring syntax checks)
All of these can be replaced with a single ruff call. Ruff consolidates the rules from all these tools into a comprehensive and non-overlapping corpus. And removes the burden of having to find the right invokation order.What's missing for ruff to be the gold standard, is to adopt features from:
- autopep8 (wraps long comments)
- docformatter (docstring auto-formatting)
- black (determinism code formatting)
- blacken-docs (applying black on Python code blocks in documentation)Though it's probably a bit conservative. Some rules are implemented as straight one-to-one ports from Pylint, and those are easy to check off, but others are Pylint rules that we've already implemented under other names, and those have to be tabulated on a case-by-case basis.
Perhaps running Ruff plus a type checker gives us close to what Pylint does today? Pylint is pretty comprehensive, and awesome for that, but I'd love to lint at the speed of Ruff.
At a previous company using the jvm, we used 6x t3.small instances at AWS to handle a much bigger load than what we now have 100+ pods handling in python.
#1 Second-mover advantage. It got to learn from all the existing tools, skip all that.
#2 Deduplicated work. Three separate linters means parsing the code three times. A “do everything” does it once.
#3 Rust is faster than Python.
These things combine into someone showing up and suddenly making a tool that a lot of people are using to replace 3+ tools. But that doesn’t mean the previous versions are bad. They were necessary to be where we’re at.
Pydocstyle, pylint, and ruff will all check for some amount of presence and format, but if a type is wrong or an arg missing, you're on your own, AFAICT.
I found pydoctest¹, but it looks somewhat unmaintained, and bugs like "doesn't work with relative imports" makes me hesitant to use it.
A lot of folks want us to add that kind of enforcement to Ruff (https://github.com/charliermarsh/ruff/issues/458), and I want to do it, but it's a big project and some other stuff has higher priority :)
Totally understood this isn't a top priority, but very glad to hear it is potentially in-scope. I've been looking for an excuse to get into rust... we'll see.
Now, we also have Rust in the performant side of things, and the guarantee that 'if compiles, it'll work' (aside from logic bugs), is quite a big one. Also, considering how helpful the compiler is, and the mentioned guarantees, more people are enthusiastic about having faster things.
In short, it shines for this problem domain in my experience.
I do think Go has have success in the same space because it is fast, but IMO the ergonomics in Rust for this problem space are a bit better. Of course, that's just my opinion.
> You can still have kinds of memory bugs in all of these incl Rust, esp leaks.
You're right. Leaks are not covered in the memory safety guarantees by Rust, neither are OOMs, but you can't deny you're on a way more safer space than with either C or C++.
For example Babashka is distributed like that (https://github.com/babashka/babashka/releases)
Just today I made a simple app in C translating text files to an array you could put in a .c-file, as a piped build tool with a small twist compared to xxd. I thought about doing it in Python first, but I realized it was way simpler to do in C. Both as a user and programmer.
I guess Rust have the same lure.