"We're building a new static type checker for Python"
twitter.com
twitter.com
Ruff truly can become one tool to rule them all.
This is true! We're also contributing salsa features upstream where we can, e.g. https://github.com/salsa-rs/salsa/pull/603
A fast Rust based type checker would be amazing!
I've personally found that dealing with mypy has been more noisy and painful than using pyright and leads to a lot of users/other devs just ignoring typing altogether.
I think that type checking has to closely match the semantics of the language and if there's a gap, it will often push users to do the easy thing, which is just ignoring checks.
To understand shortcomings of MyPy, I strongly suggest reading pyright's documentation for how they compare: https://github.com/microsoft/pyright/blob/main/docs/mypy-com...
Quoting the pertinent part:
> Pyright was designed with performance in mind. It is not unusual for pyright to be 3x to 5x faster than mypy when type checking large code bases. Some of its design decisions were motivated by this goal.
> Pyright was also designed to be used as the foundation for a Python language server. Language servers provide interactive programming features such as completion suggestions, function signature help, type information on hover, semantic-aware search, semantic-aware renaming, semantic token coloring, refactoring tools, etc. For a good user experience, these features require highly responsive type evaluation performance during interactive code modification. They also require type evaluation to work on code that is incomplete and contains syntax errors.
> To achieve these design goals, pyright is implemented as a “lazy” or “just-in-time” type evaluator. Rather than analyzing all code in a module from top to bottom, it is able to evaluate the type of an arbitrary identifier anywhere within a module. If the type of that identifier depends on the types of other expressions or symbols, pyright recursively evaluates those in turn until it has enough information to determine the type of the target identifier. By comparison, mypy uses a more traditional multi-pass architecture where semantic analysis is performed multiple times on a module from the top to the bottom until all types converge.
> Pyright implements its own parser, which recovers gracefully from syntax errors and continues parsing the remainder of the source file. By comparison, mypy uses the parser built in to the Python interpreter, and it does not support recovery after a syntax error. This also means that when you run mypy on an older version of Python, it cannot support newer language features that require grammar changes.
Astral's type checker seems to an exercise in speeding up Pyright's approach to designing a type checker, and removing the Node dependency from it.
If I am not implementing a LS, then how is it of any importance, whether the type checker was designed with typing a LS? How does that benefit me in my normal projects?
If there are no semantic improvements, that allow more type inference than MyPy allows, I don't see much going for Pyright. Sounds like a "ours is blazingly faster than the other" kind of sales pitch.
I’ll still be switching to the astral offering as soon as it’s production ready.
Like I said, this was many years ago - mypy might've gotten slower, but computers have also gotten faster, so who knows. My hunch is still that you have an issue with misconfiguration, or perhaps you're hitting a bug.
Part of the problem for me is how easily caches get invalidated. A type error somewhere will invalidate the cache of the file and anything in its dependency tree, which blows a huge hole runtime.
Checking 1 file in a big repo can take 10 seconds, or more than a minute as a result.
I found cases where it would even treat `foo: SomeType` and `foo # type: SomeType` differently!
I tried to fix that one but looking into the code lowered my impression of it even further. It's a complete mess. It's not at all a surprise that it gets so many things wrong.
Overall Mypy is like kind of like a type checker written by people who've only ever seen linters before. It checks some types but it's kind of wooly and heuristic and optional.
Pyright is a type checker written by someone who knows what they are doing. It is mostly sound, doesn't just say "eh we won't check that" half the time, and has barely any bugs.
Seriously check the closed/open issues on Github - there's a touch of "I disagree so I'm closing that" but only a touch. It mostly has so few open issues because the main author is a machine.
The only real problems with it are performance (it's ok but definitely could be better), and the slightly annoying dependence on Node.
>By comparison [to pyright], mypy uses a more traditional multi-pass architecture where semantic analysis is performed multiple times on a module from the top to the bottom until all types converge
That makes it sound like mypy does in fact do The Right Thing (albeit the slower, less-suitable-for-LSP thing), rather than a messy pile of heuristic/optional hacks.
No, but they did kinda say that about `attrs`, which is a big deal for me and my stuff at least. Hopefully this new project can deal with all the crazy dynamic stuff python allows you to get away with.
Unlikely. It's best not to write code like that in the first place.
But without that, I always felt like I was actively fighting mypy. It seemed like it was written for a totally different language than Python.
Compared to another more modern type system like TypeScript, sometimes you don't explicitly type something and yet TypeScript usually does exactly what you expect.
> Duplicate of #5613, still low priority, better use dataclasses, as suggested above.
1. https://github.com/python/mypy/issues/5944#issuecomment-4412...
https://github.com/google/pytype?tab=readme-ov-file#pytype--...
[1] You don't have to typehint everything for this to work. Pyright infers return types just fine for instance
Anyway, agreed that this is very exciting news. Poetry was great, and then I found uv. It's... wow. It's really good.
Web devs; I thought that was implied.
That said, Node is awful as a backend language, and speed has nothing to do with it. I will write Go, Python, Rust, or anything else because of how those languages have been designed. JS doesn’t belong in my backend—YMMV.
Unfortunately that massively limits its performance too since loading a large django codebase is pretty slow.
Here's the github issues filter linked in the screenshot:
https://github.com/astral-sh/ruff/labels/red-knot
And the best answer/description of what the type checker will be:
Missed opportunity to call it typy in my opinion
As you say though, not everybody gets to choose their language, and I certainly wouldn't choose Python if I could possibly help it, so if we can add good static types to Python, it's still not a great language, but it's much improved by static types.
If anyone built a typescript layer on top of Python, I’d cheer that effort. If that new thing from the article will be that, I’m stoked.
People asking why statically typed python: consider that it's the #1 language on GitHub (assume ts and js are different) and LLMs are really good at generating it.
In my case, I have, in the pursuit of my engineering goals, had to dig into a Python component that I did not normally work on or intend to work on. I am thankful it had type hints.
I don't have anything against typing. I just get annoyed when it gets out of control in a language it wasn't part of to begin with. Just look at how complex typescript typing is to see where it can go.
Used pragmatically it's fine. When it's out of control we might as well switch to another language.
Seems pretty plausible that someone can dislike python and still like a job because it checks other boxes.
I hate every language! For different reasons. But I still like programming.
> But most people in our profession get to choose their job, don't they?
Let's say someone has only worked in Python professionally but dislikes it. They spent nights & weekends learning Rust. Now they want a Rust job. Because a Rust project would not hire a Python programmer, this person who hates Python is stuck working on Python projects.
It's attitudes exactly like "Why would a Python project hire a Rust programmer?" that keep people stuck in languages they don't hate. Those attitudes are the reason a senior software engineer with over a decade in the field can't switch tech stacks. Without that attitude, job postings would be written for programmers, not python programmers, and the job market would be better for it.
Choose, yes, but from a finite set.
> Men make their own history, but they do not make it as they please; they do not make it under self-selected circumstances, but under circumstances existing already, given and transmitted from the past. The tradition of all dead generations weighs like a nightmare on the brains of the living.
— Karl Marx, "The 18th Brumaire of Louis Bonaparte"
you can obviously still use it untyped to hack together scripts, but I still find myself leaving annotations for clarity.
I've been using it since the mid 2.x days and find the style we were writing back then to be incredibly dated.
This is grossly incorrect.
Personally, I'm one of those people that got converted over to statically typed languages, as I spent a large part of my early career in startups using Ruby. I don't think I could ever go back after using Go, TypeScript, and Rust over the past decade or so.
There's also something fundamental about taking a dynamic language like Python or JavaScript and adding typing versus taking a static language and adding dynamic typing (e.g.`auto` in C#). The dynamic language allows modifications that are _really_ convenient if not a little hacky that you just can't easily express without big refactors in static codebases. Things like tagging objects with properties, making quick anonymous structs (before that became a thing in modern languages) so you can return tuples of values, other constructs like dynamic functions or whatever. Progressive typing is just so much more expressive
Assuming you meant C++ - no, auto is not a even a tiny bit dynamic. It's still fully static typing, but with type inference.
The compiler will still prevent you from calling a method which doesn't exist.
Looking at my Python before I began using type hints and data structures such as dataclasses, typevars, etc I was still type-checking a function's argument, but it was very manual and messy. Nowadays, my code is so much cleaner, concise, less error-prone, and much easier to revisit after a long period away. IMO there isn't a single downside to leaning heavily into types, even if they're sorta "fake" like they are in Python.
And full typing is also a great first step for rewriting in another language.
I used to prefer dynamically typed languages for most tasks, and saw attempts at statically typing everything as some kind of [compulsion][1].
At work, we were writing a bunch of tests and tooling in python, and the team lead at the time insisted that we type things and lint with mypy strict as if it were a compiler.
I said to my teammate "this is so stupid, if you want to know the types of the parameters, you read the documentation comment right there in the definition. If you want to know more, you read the test. If you want to know even more, you read the source."
My teammate replied, "YOU will write a function with good documentation, thorough tests, and a comprehensible definition, but as an organization grows, a shrinking percentage of code will be written that way. These automated checks prevent the worst from coders who care less than you do."
I had no retort, so there is that.
[1]: https://github.com/dgoffredo/stag?tab=readme-ov-file#why
Type hints fix this by allowing you to ensure your contracts are upheld while your team continues to move. If you don't find yourself needing them, you probably don't have a team that's trying to ship and iterate super quickly, or your project is still at the one-off script scale.
Who? People on your team or randoms complaining about your open source project? If it's the latter, you can tell them to get bent or write their own type stubs if it bothers them. Type hints are fully optional in Python.
Python typing is about machine checkable properties of your program. Consider a function foo that takes a dictionary keyed by uuids, reads from it, and returns some kind of information. We can check all sorts of properties without ever running the code. Accidentally construct a dictionary keyed by strings instead of uuids? Rejected, wrong argument type. Accidentally modify the dictionary in the body of foo()? Rejected, if you wrote Mapping[UUID, Any] for the argument type. Mapping alone does not allow you to assume mutability. So many of the errors I make in Python, I can declare my intention not to make, and then check to see if I’ve committed them.
Also its super hard to find non-crypto-currency rust jobs, but python jobs are pretty common.
But now over the last years, I happen to prefer changing programs a lot without running them a lot. I often do code refactorings over multiple commits, being highly concentrated, without running the code once. And then only running the code only 1 hour or 2 later. Such a pleasant activity! Just writing, reading, writing, thinking code. No ugly and long stack traces, no need to input user data into input fields, just reading and changing the program code, with the editor in full screen.
And guess what I highly started to appreciate, after many years of Python! Type annotations! With type annotations a LSP would immediately underline any mistake I would accidentally make, and that made me feel much more comfortable and secure about my changes, over several commits, touching 100s of lines without ever running the program. I suddenly learned to understand, after many years of dynamic type less programming with Python, that I can get more easily and for much longer into a flow state when I add annotations and run a LSP live over my code while I work on it.
Ruff is amazing, uv literally fixed all the python packaging issues we have, and now a type checker!!
Can't wait to see it!
They only try to create enough value that the VC can one day sell their stake for more than they paid for it. The key is convincing the VC not that that is particularly likely, but that if it happens there's a chance it'll be a lot more than they paid for it.
[0] occasionally a company gives money back to investors when it fails
> What I want to do is build software that vertically integrates with our open source tools, and sell that software to companies that are already using Ruff, uv, etc. Alternatives to things that companies already pay for today.
> An example of what this might look like (we may not do this, but it's helpful to have a concrete example of the strategy) would be something like an enterprise-focused private package registry. A lot of big companies use uv. We spend time talking to them. They all spend money on private package registries, and have issues with them. We could build a private registry that integrates well with uv, and sell it to those companies. [...]
i have a similar concern for the Pydantic folks commercializing via observability platform; how big is the market for python-only observability?
i think both cos will either look for money elsewhere, or find themselves having to deal with at least a few other language ecosystems. Even python+typescript would 10x the TAM or more
Essentially, VC's know very that what they're funding is risky and may go completely caput. However the VC model explicitly aims to fund companies that have the potential of becoming huge and becoming very profitable.
Both the founders and VC's behind Astral are not stupid, they know that it could fail or only have moderate success... However, if they succeed in building the next generation go-to ecosystem for python development (which is arguably the world's most popular programming language), we could imagine a future where they spin off some product and it does incredibly well!
I don't know the solution, and I really think these tools look about as great as tools can look, so I'm very keen to use them. I would just like some way to know that they'll be supported for 10 years at least.
Django has a lot of magic that makes type checking more difficult than it should be.
I'm not really expecting anything, even just the speed bump will be nice.
Moving to a Pydantic-style model would be a whole different ballgame. Hell of a lot of work, and meanwhile we can’t even really get to “let’s have type hints in some of the Django infernals”.
But! I do think someone who wants to have fun could try to build a type hint based Model subclass as a third party lib! There’s a looooot of stuff to resolve in that world but for any proposal like that the general vibe is “prove the feasibility with a third party package first”.
It's essentially what you described and it's growing in popularity.
At the end of the day I end up writing a service layer over everything so I'm not super bugged by most of this. But it's busywork when first starting up a project.
https://adsharma.github.io/fquery-meets-sqlmodel/
https://github.com/adsharma/fastapi-shopping/blob/main/model...
All type checkers other than mypy (e.g. pyright, intellij) have ignored the level of plugin support necessary to make django work well, and so they are DOA for any large existing django codebase. Unless ruff decides to support such a dynamic interface as mypy's, it'll fare no better.
We use mypy with [django-stubs](https://github.com/typeddjango/django-stubs) which works ok nowadays.
There was an effort to create a typechecking plugin interface for dataclass-style transforms for python type checkers, but what was merged was so lacking that one couldn't even make something close to django-stubs with it.
Tangential, I'm waiting for Edgedb to become next SQL so all ORM moats like Django's are drained.
At no point.
What probably will happen is that both the JS and the Python reference implementations will have first class support for type systems which will allow them to generate optimized native code for the typed sections of your code while keeping the untyped sections as bytecode or as higher level compiled code (i.e. a bunch of runtime function calls dealing with tagged union objects).
if you would like to read more on the topic the keyword is "gradual typing"; in particular you might be interested in looking up stanza, a language designed around gradual types from the outset rather than retrofitting them on top of a dynamically typed language.
what are places to discuss gradual typing these days ? I stopped looking a few years ago (after clojure and scheme started projects on that topic)
Javascript has a uniquely privileged position as the native browser language, and a couple of decades of effort has been expended on making it usable.
Everything Astral have put out has been a step improvement to my python dev flow. Until they disappoint me, I'm staying on this hype train.
I've not checked it out enough yet, and it's early days for it anyway.
Just saying it may be worth keeping an eye on.
I've seen some videos of Chris Lattner talking about it.
curl -ssL https://magic.modular.com/deb11637-c8c3-4f76-9195-85af25ef2e... | bash
I have been seeing this pattern for some years now, and also seen criticism of it.
do they not care about security, or is it not a security issue? conditionally or unconditionally, that is?
We have dialyzer, which is very polarizing and frequently maligned, and somewhere between 2.0 and 3.5 partial implementations of LSP depending on how you're counting. It's been radio silence ever since the last public blog post announcing the formation of the new LSP team/project and the incumbents still ship new code periodically, so I have no idea where that stands. None of them hold a candle to r-a in their current forms.
Credo is pretty solid, and Adobe's styler is neat but I wouldnt blame anyone for sticking to vanilla `mix format`. ExUnit is also fine. I'll never touch Ash or Igniter, as they don't speak to my personal or professional values.
I can't speak to the type-system stuff in 1.18, as I have a bunch of "code-complete" dependencies that won't build without compiler warnings. Outside of that blind spot I consider Elixir to be in an ongoing tooling slump, personally. Certainly when compared to more innovative and populous ecosystems.
(PyContracts and iContract do runtime type checking, but it's not very performant.)
That MyPy isn't usable at runtime causes lots of re-work.
icontract: https://github.com/Parquery/icontract
The DbC Design-by-Contract patterns supported by icontract probably have code quality returns beyond saving work.
Safety critical coding guidelines specify that there must be runtime type and value checks at the top of every function.
It isn't much faster though so hopefully this will fix that.
> The entire system is designed to be highly incremental so that it can eventually power a language server (e.g., only re-analyze affected files on code change).
- why not just contribute to the community tool?
- there's already a major split in Python type-checking tools, if there's a third that doesn't agree with either of them it'll be a mess for projects to deal with
- astral has been hiring like mad recently and has yet to communicate that they can actually make money ($5 million doesn't last forever)
- does it actually exist? is this currently a closed-source codebase, or is "we're building" future tense?
From the thread:
> We haven't publicized it to-date, but all of this work has been happening in the open, in the Ruff repository.
- re: the money: https://news.ycombinator.com/item?id=42869358
what community tool? mypy is written in python, and is non-incremental, whereas astral's goal is to build a fast, incremental type checker in rust. there is no way to get from one to the other via code contributions, they fundamentally need to start from scratch with their own architecture.
as for the split in type checking tools, it is not as bad as you think; the syntax and to a large extent the semantics of the type system are defined via a community process and standardised upon, and the various type checkers largely differ in their implementation details but not in their interpretation of the code. so you can freely use several different type checkers without fear of disagreement.
those people producing fantastic tools that have transformed the ecosystem of linting and packaging should be working on it instead of working on transformative tech
I'm surprised you already know it's not as bad as I think - have you been able to use it? Working on a team that mixes mypy and pyright is pretty frustrating since they don't agree on everything (i.e. when one changeset passes on one and fails on the other) and I see no reason to believe the inconsistencies will become more rare when the number of opinions goes from two to three
also if you do find a case where two type checkers genuinely conflict on a piece of code, one or both of them would definitely like to see it as a bug report. if the underlying cause turns out to be undefined behaviour in the specs, the general typing community will work together to nail it down and the type checkers will all adapt. in general it is a very cooperative process that values the existence of a common set of standards, which is why I think having more type checkers will only improve the situation wrt nailing down corner cases in the specs.
I'm happy you believe that the community can converge on unified standards - I wish I shared the optimism - but that's not where we're currently at and years-long efforts don't help me now. (For example, I'd love to use PEP 695 generics since I found a use case for them around 18 months ago but I can't until I can get away with not supporting 3.11, which is years out.) Maybe everything is perfect in 2028 or so - I'd be thrilled - but that doesn't help the pitch for annotations being a value add for people who are worried about the current jobs.
pep695 is a good example of how the spec is still evolving to try and make the explicit type system work better for python developers. dataclass transforms are perhaps an even better example - that's a pattern that is very specific to python and based on real world, non-typing-related ways in which people use the language, and which they wanted the annotation system to cover, and there was a collaborative effort to develop a way to do it.
also note that there are escape hatches like error suppression (e.g. if pyright likes some code but mypy wrongly thinks it's a type error, you can add a comment so that mypy ignores that one error while continuing to have pyright check it) and typing.cast (which is a directive all the type checkers honour, to say "trust me, this variable has this particular type whether or not you can verify it statically").
another thing to consider is that having type checking work interactively in the IDE is something a lot of people want, and pyright was developed specifically to support that feature. mypy, pytype, and pyre were all architected as batch type checkers, i.e. they need to run on a complete program as a standalone process, and it would have been anywhere from difficult to impossible to get them to support the type of incremental checking an IDE requires.