Pyre: A performant type-checker for Python 3
pyre-check.org
pyre-check.org
I tried mypy [1] before but it was a bit cumbersome to keep it checking my code. At work I use pytype [2]. It's good enough but that's only because someone else made a build system integration for me. Pyre's scan-all-the-files approach seems easier to get started. Is there any catch? Which one are you using and what's your impression?
[1] https://mypy.readthedocs.io/en/stable/
[2] https://github.com/google/pytypeOr does it mean that we should check types using all type-checkers available to have the most compliant code possible?
I think for most people any of the 4 big ones is fine. Most of the common type errors will look identical regardless of which you pick. Configuring multiple in CI is pretty straightforward if you want to go full. My current codebase has mypy + pyright configured. Pyright was a mix of fast response time on github issues + I use vscode and was the one that added type checks to CI.
If you want a deep dive pycon is this week I think and there's a day for type checking related talks including one talk comparing the 4. I think type checking day is thursday.
Hope they post the recordings somewhere later.
Pyright's VSCode integration, Pylance, provides a great auto complete experience. So I use Pyright for autocomplete in VSCode and mypy for type checking.
Worse, importing third party packages often fails silently and when you can get error messages they tend to be completely inactionable (I recall one error message which linked to a web page that had lots of details and workarounds for fixing other problems, but none of which solved the error itself).
Figuring out how to distribute my own type annotations was similarly painful. IIRC, you have to drop a specially-named file into a particular directory and this is all undocumented save for a dense PEP and the error messages are unsurprisingly terrible. All of these things are actionable, but the progress seems slow (these have been among my top grievances since the project debuted, so the maintainers and I have different priorities, clearly).
The problems for which I'm less optimistic tend to revolve around shoehorning typing into existing Python syntax--e.g., to get a callback that takes kwargs you have to define a protocol with a `__call__` method that takes kwargs because you can't express it with `typing.Callable`. Similarly where a language with first-class support for types might have `type Foo<T>`, Python makes you write `T = TypeVar("T"); class Foo(Generic[T])` or something like that, and it gets more confusing when you only want one of the methods to be generic and I can never remember whether that `T` takes on a single type across all uses or which scope I need to define it in, etc. This is largely an ergonomic nightmare and I don't have lots of optimism for this stuff to improve unless the Python community really comes to embrace typing as the default way to use Python (but I suspect most people who care a lot about this kind of stuff will leave for Go or other languages where these things just work out of the box).
I don't know about you but I find PEP 484 very clear here: https://www.python.org/dev/peps/pep-0484/#scoping-rules-for-...
The nice thing is that, since TypeVars are ordinary variables, you can re-use them / import them in multiple modules and avoid a lot of boilerplate code.
Maybe I'm missing something, but the boilerplate only exists because Python makes you define them as ordinary variables in the first place. So while you can have a `foo.py` file like this:
T = TypeVar("T")
And a `bar.py` file like this: from foo import T
class Bar(Generic[T]):
...
So are you saying that this is less boilerplate than a `bar.py` that just defines its own `T` TypeVar? Because that seems like a pretty comparable amount of boilerplate. Or are you saying that it's less boilerplate than languages that don't treat TypeVars as ordinary variables, e.g., Rust? Because in Rust our `bar.rs` file would look like this: `struct Bar<T> {...}` (and the scoping rules are patently obvious, to boot).You might be interested in the discussion over here then -> https://github.com/python/typing/issues/769#issuecomment-741...
Lack of recursive types has been a major deal breaker for me. If you have a class Foo that can construct a class Bar, and a class Bar that can construct a class Foo, you can't express that in mypy without Any-Generics.
To me, at this point, mypy is barely a linter, it's more of a "this helps my IDE autocomplete faster", and definitely not a type system.
from __future__ import annotations
class X:
def __init__(self, value: int) -> None:
self.value = value
def to_y(self) -> Y:
return Y(self.value)
class Y:
def __init__(self, value: int) -> None:
self.value = value
def to_x(self) -> X:
return X(self.value)
x = X(10).to_y().to_x()
print(x.value)And by "successfully" I mean "it ran at all". Every other thing I've thrown at it has caused it to crash. I quite like the concept, but implementation has been rather abysmal as far as I've seen. Has it improved in ~ the past year?
I'm planning to migrate this to my own blog soon.
It does seems to support the entire-directory parsing as well. That's nice! Let me give it a try next time. Thanks for the tip!
As a reminder to myself, here is the link to the doc: https://google.github.io/pytype/
It's a great tool when editing code. But you still should run Mypy or another checker for more-correct analysis.
We really need a modern Python alternative. I don't know of any that don't give up the REPL / single file script features which are pretty huge advantages of Python to be honest.
But I do agree Typescript is one of the best alternatives today if you can stomach `tsconfig.json` `eslintrc` and `node_modules`.
Then there is Julia as well.
1-based indexing is definitely not a flaw. And version 1.6 massively decreased package import and precompilation time.
It really is. We've known it for literally decades:
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
My own preference is random indexing, it's so much more exciting to be surprised https://www.juliabloggers.com/random-based-indexing-for-arra...
I do agree in some applications like data science / machine learning you could never get people to switch from Python because everyone uses it. But Python is also used for loads of other things, e.g. hacky build systems, web scraping, etc. that could easily switch.
Guile has many useful tools, like OS level threads and also a fibers library, community projects, a good manual, albeit sometimes lacking a few examples, an active community and mailing list and more. Since it is a Scheme, it adheres to a Scheme standard, which specifies many things already. Then it implements many SRFIs, which also specify many things. I guess, that you name something that goes beyond the Scheme standard and beyond the SRFIs "weird".
Could you point out what more specifically you personally find weird about it?
That's what's weird.
Usually when someone claims, that something is "weird", I want to know, why they think so. When it is about programming languages, I would like to know what it is exactly, that they think is weird about that specific language and that is what my question entailed.
Then, macros and DSLs are awesome for solo cowboy coders writing code, not so great for professional programmers working in large teams and reading code 10-100 times more than they write code. This also leads to fragmentation and half finished solutions since the solo devs generally scratch the itch but don't do the hard work required by the last 20% of the project (which as we know, takes 80% of the time).
Then, adoption. It's not there. There are no IDEs except for Emacs (not a popular editor/IDE) or commercial ones which are super expensive. Libraries are in much lower quantity and variety and frequently not as good as those of mainstream languages. Etc, etc.
Macros, DSLs, well, of course you can abuse then, like anything else in computer programming. However, there are many examples of how they can be used in a great way. Look at some Racket macro things like typed Racket for example. Or look at pipelining operators. Or timing. Or memoization. All these are very well usable and there is no problem with using them in a team. Well written macros allow taking cool features from other languages to your Scheme dialect of choice.
Emacs is still well liked. I recommend you get on the mailing list and read a few weeks about how varied its usage is. Very active mailing list.
Libraries of lower quality? Even "much lower"? Where is your source for that? Not sure which specific ecosystem you have looked at, but that experience is completely different from mine.
Aside from the fact, that I can usually solve the problems by just using Guile features, Scheme and SRFIs, not even needing an external library, there are very clever people active in the ecosystems of lispy languages (including GNU Guile) and FP languages, outputting high quality code, often going beyond what some mainstream language library does, while using good abstractions to do so.
It is important, that languages like GNU Guile, which implement interesting and powerful concepts, continue to attract people, who want to learn more than the mainstream fad and improve the status quo. It is a great journey of learning, which I recommend to any software developer looking to widen their horizon and to improve their skill.
Conda is so slow that I sometimes wonder if we are being trolled by some cruel God of programming. Pip is faster, but version resolution is iffy.
Hell, even just assigning versions to python packages is nothing short of ridiculous. Do you use version.txt in the root folder and set it manually? Do you have it set from SCM? Which of the half a dozen packages do you use to have it set from SCM? setuptools_scm? Versioneer?
There are a set of tools in the python ecosystem that have basically no equal in any other language and these tools and the surrounding mindshare make python irreplaceable in the near term. The language itself is easy to learn and powerful enough to be able to do data analysis with ease. Good python code is easy on the eyes, which I personally consider an important aspect.
Outside of these tools, core parts of the ecosystem are basically an XKCD joke.
I tried switching to Julia as I find both the language and the ecosystem are vastly superior in their foundations. Unfortunately the maturity is not there yet, and neither is the mindshare. If I had to bet my career on adopting the language in a business setting, I'd not be prepared to do so. Which is a shame, because the situation turns into a Catch-22.
What tools do you find irreplaceable?
All the tools you listed do different things, except maybe poetry and pipenv, so you can pick whichever one you like, you're not supposed to do anything. You can have choice, illusion of free will, etc...
As to module version, there is a standard on how to define it in __version__: https://www.python.org/dev/peps/pep-0008/#module-level-dunde...
Strongly disagree. Conda, pipenv, pyenv, venv, poetry are all trying to solve the same problem (although conda tries to solve some other problems too).
Choice is not always good, this is why we have standards.
I would recommend pyenv. I understand why people are attracted to poetry, but pyenv arguably offers all of the same benefits that poetry has as well, and has better adoption.
Not really, some of these manage the issue of multiple system pythons (pyenv), while some manage isolated envs for particular projects (venv), and some try to be wholistic python project and dependency managers (pipenv, poetry, arguably venv + pip freeze, but that's not "wholistic"). Conda sort of tries to be all of the above as well as a bunch of other things (high performance options etc.)
There is a ton of overlap, denying that is disingenuous.
I use both pyenv and poetry, and basically nothing else apart from the occasional call to pip.
I cannot imagine dropping poetry and only use pyenv, but I can imagine dropping pyenv (have been looking at asdf recently...)
If I were pip dictator, I would try to make pip the one tool to handle all python packaging and environment management. In particular, that means pip would handle the management of different Python versions, different Python environments, native dependencies, running tests, making builds perfectly reproducible, releasing new versions of libraries, and creating new projects.
"Creating new projects" seems like it is not a big deal, but in practice I think that if there were simply a "pip new" command that set up a new project using the best practices advocated by the pip team, it would go a long way toward standardizing the ecosystem here.
The first time you have to run it as:
conda install -c conda-forge mamba
From then on, you replace conda with mamba. For example, if you are installing dask-cuda from the rapidsai channel you run it as: mamba install -c rapidsai dask-cuda
At this point, mamba is just so much better and faster, that it's the first package I install in an Anaconda environment.Mypy and Pyre are pretty similar, but Pytype has very different standards.
In C, you ultimately care about what the compiler says. And this has also led to dialect-specific C code that works fine in one compiler but doesn't compile or runs incorrectly in another compiler.
If you are using the most recent pep features possible than yeah you might have an issue. That's similar to clang/gcc both taking time to implement new c++ standards and not being compatible there. If you only use pep 484 which covers most basics well then you should be good for any checker.
This is totally untrue? Pyre adds completely new notions like taint analysis, pytype adds a totally different inference system - the type systems are not compatible and it has nothing to do with implementation of PEPs.
The difference with C is that all of the other linters and analyzers are on top of the base type system, they don't replace it. In the case of Python the only standard part is where the annotations go and how they get resolved at runtime, it's otherwise completely up to the type system implementation to determine what's what, and all 3 of them do so in very different ways.
This is like if clang/gcc had totally different behaviors for `auto` in C++, but really it's not even comparable.
My base statement of 4 main type checkers are pep 484 compliant and that's main type checking most people refer to for python. If you get a type error that's inconsistent with that pep that's a bug. If you use a feature outside of any pep like plugins, taint analysis, pyright stub generation, then yes that is unspecified and can vary. Most engineers I work with python type checking stops mostly at basic type system and I'll sometimes point them to newer peps for better type definitions. My actual experience is most places don't consistently do any type checking at all although people tend to be open to it if you're fine guiding them/setting it up.
$ python -c 'import this' | sed -n 15p
There should be one-- and preferably only one --obvious way to do it.
Sadly, the python type-checking situation appears to be going the way of the python packaging situation.Issues seem to be from libraries with crappy, incomplete, or incorrect typing/stubs, not from the checker itself.
Every day I tire more and more of using Python for larger projects for these reasons
https://www.amazon.com/dp/B01MRJVUPC/
Maybe not the best science fiction book of all times, but it left a mark on my memory... and curious things like this name have stayed there since when I was a kid avidly reading everything that fell on his hands :-)
-- EDIT: Link to the book
I personally think it's mostly a waste of time to do type hinting for the purpose of catching errors in a strongly typed language like Python. Type errors just aren't practically a problem in dynamic languages. Doing types for performance is a great reason to do it, though.
For starters, type checking can't actually guarantee you won't have type errors at run time. Because, unlike in a statically typed language, in Python, anyone can always choose to just not use type hinting. Whenever that happens, as far as the type checker is concerned, anything goes.
But type hints are very useful as hints. They help with editor tooling, which can make it easier to navigate an unfamiliar codebase. They provide extra information that makes the code easier to read. And they give me an opt-in form of type linting that allows me to set up regions of code where I don't have to take quite so much personal responsibility for ensuring that arguments are compatible with parameters. In short, it's not a correctness prover; it's an energy saver.
For large, multi-developer projects, it is impossible to exaggerate the benefits of static typing.
Typecheckers like mypy are (1) optional, and (2) are only as good as the type hints themselves. Unlike with a language that's actually statically typed, the relation of the type annotations to the actual types of the data operates largely on the honor system, so the type checking is more a validation of internal consistency than full-on static type safety. Which is fine, and well within Python's "We're all adults here" ethos. It's just that it's still not anywhere near the level of rigor you get out of the type checker in a language like OCaml.
I disagree that type errors "aren't practically a problem". NoneType errors and AttributeErrors turn up all the time in Python.
I've also written about this publicly before, e.g. https://mail.python.org/archives/list/typing-sig@python.org/..., where I found a number of errors by tweaking pytype in a way to make it stronger. It's difficult to quantify what percentage of bugs would have been found anyway, but pretty much any time the type systems I work with improve, it uncovers new bugs.
The result might be surprising, but those are the most interesting results! I note that, so far, nobody has countered this with any evidence, only anecdotes and "surely not!"
I'm not a fan of this kind of methodology, because there are so many different use cases for github issue tracker (including feature requests, documentation requests, etc.) that I think the denominator here is overestimated. I'd rather have a study that looked at 1000 issues in depth with a rigorous rubric than one that grepped through millions of issues.
Also, looking only at the issues doesn't count the ones that were caught in development or testing. I like the notion of "moving your bugs to the left"[1] and would much rather find these type errors at build time than later.
Finally as I said in my other comment, avoiding type errors is not the only benefit of type checking. Making code easier to understand and refactor (including programmatic refactoring) are other benefits that are not captured only by counting bugs.
Aside: It's worth commenting more on "code readability" because it cuts both ways. It can absolutely be the case that a complex, manifest type system can be a hindrance to understandably and productivity due to the "noise" of all the types that need to be written. I think struggling with C++ and (to a lesser extent) Java typing is part of the reason the industry really embraced dynamic languages in the 2000s. Python and Ruby were like a breath of fresh air. But eventually we built large systems with these languages and discovered the pain that comes from having no type-safety.
The trend I've seen in the last decade is to embrace type inference, with explicit type declarations on function signatures and otherwise only when needed for disambiguation. To me, this is the sweet spot, giving the benefits of static typing with much of the cleanliness and concision of dynamic languages.