Ruff – a fast Python Linter written in Rust
github.com
github.com
The linters the author is comparing against have significantly more features an more overhead given they need to support many more complex rules/situations, also I'd argue this linter is far from well organized given it seems all checks appear to be implemented in a flat file [1].
I would not be surprised if a fully fledged linter written in rust outperforms any linter written in python by a factor of 20-25x though.
I'd be curious to see functionality & performance comparisions if this project continues.
[1] https://github.com/charliermarsh/ruff/blob/0b9e3f8b472dc3fd0...
I've tried to be honest about ruff's limitations: it supports only a small set of rules, it's not extensible in any way, it's missing edge-case handling, it's significantly under-tested compared to existing tools, etc. I don't consider it production-ready -- it's a proof-of-concept.
(I _did_ try to build conviction that there weren't inherent reasons for ruff to slow down significantly as the set of checks extended, but I could definitely be proven wrong as the project grows...)
My goal with ruff was partly to build a fast linter, but that was really in service of testing a broader hypothesis I have about Python developer tools could evolve. If that's interesting to you, I wrote about it a bit more here: https://notes.crmarsh.com/python-tooling-could-be-much-much-.... I'm certainly not here to say that the existing tools are bad or you shouldn't use them or that you should use ruff instead. I use those tools myself -- a lot!
With regards to code organization: oh well, this doesn't bother me given the state of the project. The code will evolve as the project grows in scope, I didn't see a need to over-abstract. I'd written a small amount of Rust prior to ruff, but much of it was a learning experience for me. (Funnily enough: not that I want ruff to be organized as such, but pyflakes, pycodestyle, and autoflake are also effectively single-file projects :))
> Whereas before one might have thought that a language's tooling should be written in said language, there might come a time when that's too slow, and we instead need to move to faster languages.
> That's what has happened to Python (numpy, pandas, scipy etc are all written in C and simply provide an interface in Python), and now it's happening to Javascript as well, with Deno, swc in Rust, Bun in Zig, esbuild in Go, and so on.
On the other hand, if Python had been really committed to dogfooding at the level of something like golang, then the primary implementation would look a lot more like pypy than it would like cpython.
If you look at JS tooling like eslint, tslint, uglifyjs, tsc, webpack, ~~esbuild~~, etc, it's all JavaScript (moving to typescript if anything).
EDIT: I stand corrected in that esbuild is not written in Javascript, but I think that's a disadvantage to it's success if anything and would point out webpack is more popular
Webpack exists since 2012, of course it's more popular. For a long time there were no other alternatives, so people had to deal with its madness.
I can definitely see why libraries should be written in faster languages, they are used all over the place. And anyway, something Numpy is most importantly an interface to pre-existing high performance libraries (BLAS, LAPACK) which were already written in faster languages.
I think it is less obvious that a tooling in the second sense should be written in a lower level language. There's the obvious flexibility tradeoff, the fact that a person interested in a linter for a language probably is most familiar with that language.
I'm not familiar with the computational problems in linters. Is getting better CPU performance a huge concern?
Same thing with Webpack vs esbuild or swc, the latter two are simply instantaneous while Webpack can take minutes, even, to build an app. That it's in JS does not matter to me since I don't care about Webpack source code, I only care about what the tool itself does.
Is the standard library included with that "toolchain" tooling, or a library?
All of the libraries are linked to the project by some tooling, and a particular distribution of software could include any number of libraries (e.g., conda), I don't see why the standard library needs a special cutout.
It will get even more blurry if you start asking about runtimes, macros, and all-template libraries.
In C89, those are float.h, limits.h, iso646.h, stdarg.h, and stddef.h.
C99 added stdbool.h, and stdint.h.
C11 added stdalign.h, and stdnoreturn.h.
So those headers are part of the standard library, but they're also usually part of the compiler. The line gets very blurry.
This shows that it's entirely possible to write fast tooling in dynamic language if you know how to.
A lesson learned 20 years ago, before everything .com went bust.
0: https://programming-language-benchmarks.vercel.app/rust-vs-j...
Benchmarks are very constrained, repetitive and often numerical code that is executed a lot of times with similar input. This gives the JIT a lot of opportunity to optimise and specialise for specific types.
These tools are usually not long-running , so code has to be JITed all over again on each invocation. A lot of code will probably only be executed a few times, so might only be executed by the interpreter or low tier JIT.
There will also be a lot of AST construction, walking and modification, which is really difficult to optimise without the guidance of types. A lower level language can use better data structures that don't require so much pointer chasing. Also lots of potential GC pressure if there are multiple internal representations to construct.
JS can be insanely fast for what it is, but alot of the weaknesses are apparent with short-running tooling. There is also a reason why most complex,JS heavy web apps seem to get laggy and slow eventually...
It:
- seamlessly supports everything from old-ass Python 2 code up to very recent (like 3.10 or someting)
- it can and does exploit `mypy`-style annotations, but needs none to deeply analyze your code and actually find bugs
- it seems to do fairly serious deep PLT magicks. i admittedly haven't dived deep here but it's doing stuff that looks like escape analysis-level whole-program analysis
- it heavily caches and optimizes it's (understandably) heavy lifting via Ninja
It demolishes everything else I've tried. I will definitely take `Ruff` for a spin, we've got a lot of Python and I'm always up for discovering there's a better tool, but the dark magic PLT wizards at Google have probably gotten pretty close to what can be done.
it's worth exploring some of the other type checkers as well, since they make different tradeoffs - in particular, microsoft's pyright[2] (written in typescript!) can run incrementally within vscode, and tends to add new and experimentally proposed typing PEPs faster than we do.
[1] https://github.com/google/pytype/blob/main/docs/developers/i...
Thanks for all your work on pytype!
Specifically: analysis is broken up into two phases:
1. one that only depends on the content of a file, and 2. a second one which deals with inter-file dependencies.
The advantage is that you can very effectively cache the results of (1). This makes that phase of the analysis not performance critical. Phase 2 has to be fast, but you can push as much expense into phase 1 as possible.
This provides scalability - both at the level of GitHub where people may view code from many different commits, but also for the live editing case where you may be editing a small number of files in a much larger project. The effects of your small changes may affect many of your dependencies and you want to know the lint errors immediately.
[1]: https://github.blog/2021-12-09-introducing-stack-graphs/
I wonder if the Salsa maintainers have considered this.
And you never will. From https://github.com/psf/black:
Black is the uncompromising Python code formatter. By using it, you agree to cede control over minutiae of hand-formatting. In return, Black gives you speed, determinism, and freedom from pycodestyle nagging about formatting. You will save time and mental energy for more important matters.
Does prettier work on python code?
Tooling and libraries for Python have always been coded in compiled languages, anywhere performance matters. In this Python might differ from many other interpreted languages.
Did anybody have even a hint of doubt that a compiled linter would be at least tens of times faster than one in Python? (Either Python has got faster in recent years, or it really ought to have been hundreds. Or maybe the others rely on compiled libraries for their own heavy lifting.)