Applying Mypy to real-world projects
calpaterson.com
calpaterson.com
If MyPy isn't running in your build/CI, it's possibly worse than useless.
Until recently typed-annotated Python code wasn't checked in the build. The only time you'd be able to notice a problem was when the IDE's MyPy plugin showed a red squiggly.
One I got MyPy integrated with Bazel we could run it over our codebase, and lo and behold there were at least a dozen errors which were actively misleading developers about the code they were reading.
If something (MyPy, Pyre) isn't checking the type-annotations all the time, they're going to decay into something worse than untyped Python.
We implemented it progressively. At first I added it as a make target but didn't make it mandatory in CI so I could learn how to use it. Then I made it mandatory for a few files that I was the only active contributor to. Then I slowly added more and more files across the project, sometimes as I touched them for other reason and other times as independent changes. Eventually as mypy caught more and more bugs in other contributor's changes they started getting on board and adding type hints as well, until the vast majority of the project was hinted (we'll be getting to 100% within a few weeks).
We also use flake8, shellcheck, yamllint and black in our automated tests, as well as a couple of custom scripts (e.g. one that makes sure if you change the documentation you also remembered to regenerate and commit an updated index)
We also still write our tests to assume no type checking in parts dealing with external data, because we can still get badly typed data at runtime. But we also use valdation code to assert those types at the edges of the system so internal modules can assume types are correct.
So far, we've had a pretty positive experience with mypy and it's helped prevent a few bugs. Do be prepared to find weird edges you can cut yourself on along the way, but all-in-all I think it's a positive value add.
The first result in Google for this error points here https://github.com/python/mypy/issues/3905 which explains the issue. I'm ashamed to say I probably just read the first reply, added it to my config and carried on with my day without actually reading what it did.
`zope.interface` is more explicit and scalable than `typing.Protocol`s, and more flexible than `abc.ABC`. There's a mypy plugin for it: https://github.com/Shoobx/mypy-zope
> The drawback is that code that changes the representation of its data a lot tends not to be fast code.
That's not a very convincing reason to avoid dataclasses except in the most performance-constrained environments -- and even then I'm doubtful it'd help. Especially with `slots=True`, dataclasses can take less resources.
`slots=True` doesn't exist, though there's an experimental implementation at https://github.com/ericvsmith/dataclasses/blob/master/datacl... .
The main issue is that slots have to exist at class definition, and @dataclass takes an already defined class, so you'd have to create a clone of the class, which is nasty.
On top of that you can't always add __slots__ manually, because it conflicts with default field values (and field()s). __slots__ wants to add a descriptor as a class attribute but it can't do that if the default value is already an attribute.
I'm glad you liked the article :) Interesting tip on zope.interface - hadn't heard of that and will be looking into it
> > The drawback is that code that changes the representation of its data a lot tends not to be fast code.
> That's not a very convincing reason to avoid dataclasses except in the most performance-constrained environments -- and even then I'm doubtful it'd help. Especially with `slots=True`, dataclasses can take less resources.
re this - slotted (data)classes are useful when you want to reduce peak memory usage but don't help with the problem of programs that instantiate (and teardown) objects excessively.
IME this is a common pitfall that slows down Python programs (often by a couple of orders of magnitude) and contributes to an unjustified perception that Python is "too slow". The "excessive changes in data representation" anti-pattern is slow in every language but because the syntactic/correctness overhead of creating a mass of objects is higher in (eg) C than Python the arises much less often. My worry is that as the syntactic burden of this is lowered even further it becomes even more common.
The problem is so severe in the Java community that it has had a huge influence on JVM design. I haven't kept up with every development in the JVM but last I was using it (maybe 6 years ago) it already had a huge number of techniques to reduce the burden of frequent object creation. CPython of course doesn't have any of these tricks.
Also, does anybody know of any benchmarks of typical programs in Mypy vs cPython?
Compiled mypy is a stated potential goal - but outside the scope of this project itself. Read more here: https://mypy.readthedocs.io/en/stable/faq.html#will-static-t...
More practically, I’ve found that it’s much easier to translate well typed code to cython. It still requires work, and ymmv.
(I know they're working on it, but I seem to hit this on nearly every project, would be the number one missing feature ATM IMHO)
See this interesting blog post on how it was done: https://blog.zulip.org/2016/10/13/static-types-in-python-oh-...
Indeed, loved that closing sentence.