The different uses of Python type hints
lukeplant.me.uk
lukeplant.me.uk
I couldn't do unit tests all at once because of the design not having any way of accommodating them. I needed something to stop the enormous waste of time.
So I added type hints. The IDE should show me when illogical things were being done with parameters to methods/functions. It was fairly quick compared to a total refactoring but not effort free. I barely noticed any effect. Didn't catch a single error. Eventually I created a kind of dummy version of Android that "built" in a few seconds and I tested against that first. That allowed me to speed up changes and get some refactoring done to make a few critical unit tests and the whole thing started to get under control.
This anecdote has almost no meaning - you cannot conclude that type hints have no benefit because of one case - I just think that tests are almost always more important and hinting and the whole rigmarole of strong typing are much less of a panacea than tests are.
I've been told that Pyright is better but I've not tried it out properly. But yeah, your experience largely matches with mine, for Python at least.
The only real use-case that is both possible and worthwhile I've found is being able to say a value is a T if there's a default and an Optional[T] otherwise.
Could you show some examples?
You just need those libraries to embrace it really, then you could theoretically have type constructors that provide well-typed NxM matrix types or whatever, allowing you to enforce that [[1,2],[3,4]] is an instance of matrix_t(2, 2).
I don't see how python could possibly make such inferences for arbitrary libraries.
1. What's the datatype of the array elements
2. Does this array alias other memory
3. Is the access pattern I want to do contiguous in memory
4. If you track the provenance of an array, does it include something like a "width" dimension and a "height" dimension
5. How many dimensions are there
6. What's a good semantic description (type) for each dimension
7. As an exact integer (or modulo some power of 2 or whatever), how big is each dimension
And on and on and on. An honest-to-goodness type hint capturing that sort of crap in a way that's statically analyzable is a nightmare, and it wouldn't be totally trivial to even write the code to make a type hint like that reasonable to read and write. Even if you could, it'd probably generate a lot of noise that for any particular use of an array would distract you from the aspects you care about.
A nice hybrid solution IMO is found in that Python allows arbitrary objects to be used as type hints, and a string description of the aspects you're using/providing on a particular array works as decent documentation for other developers. For a few examples:
1. The array should describe a typical 24-bit 3-channel image. You might use a type hint like 'u8:(w,h,3)' to indicate that it's a 0-255 integer field rather than a 0-1 float field, which dimensions have width/height/channels, and that it's a 3-channel image. It'd probably be good to also label those channels with a convention like 'u8:(w,h,(rgb))', like 'u8:(w,h,3):rgb', with hungarian typing, or something (no particular recommendations on my end since I'm not usually working with heterogeneous data like that, but choosing the wrong encoding or even the wrong coordinate space for RGB or whatever is a big deal, so you'd probably want to represent that somehow).
2. You have a function signature with multiple inputs, and the computation is mostly arbitrary, but it's important some dimensions align. Then label them the same. Something like matmul(left: '(n, d)', right: '(d, k)') -> '(n, k)'.
3. You're doing some ML thing on some time-series medical data, and it's common t' have giant dense tensors floating around. Label semantically what all the dimensions are with a type like `(batch, r, a, s, channel, t)', or using longer names as appropriate depending on your audience and the background knowledge you can assume.
Libraries like einops and functions like `np.einsum` take that a step further and require stringified descriptions of the operation you're trying to do. They can have a learning curve, but the crux of the idea is that instead of writing garbage like `arr[3,6,-4:,np.newaxis,...].T.reshape(4 n, -1)` or God-forbid some sort of roll/transpose logic, you have a higher-level description.
A couple examples with einsum:
1. The dot product of v and w is `np.dot(v, w)` or `np.sum(v w) # imagine there's an asterisk; HN's parser is smarter than me`, and it's also `np.einsum('d,d', v, w)`. Arguably einsum is a bit of syntactic noise for such a simple example, but if v and w have different shapes than you think then the simpler solutions will silently produce garbage (e.g., the first option will do matrix multiplications sometimes, and the second is arguably closer to correct most of the time, but if you think you're operating on 1D objects and actually do want a channeled operation like matrix multiplication when the input isn't 1D then the sum of products is wrong and not captured in the type system), but einsum will just barf if the stated dimensions don't match your expectations. Moreover, with optimize=True it'll actually fall back to whichever of the simpler solutions is fastest.
2. Imagine you have a matrix A of shape (n, n) and a matrix X of shape (n, d) and want to compute something like A @ v @ A.T for each column v of X. You can write it via standard numpy operators, but it looks like garbage and kind of hides what's actually happening. The einsum solution is just `np.einsum('vw,wd,nv->nd', A, X, A)`. You're contracting over `v` and `w` and left with `n` and `d`. It's not perfect since you just get single-letter names to work with, but it's a hell of a lot better than equivalent options, and much easier to make suitably fast (just pass optimize=True).
And then einops is even better because roll/transpose logic is incredibly fiddly and prone to off-by-one errors in your choice of dimension or needing to deeply understand how the function works to not make footgun-style mistakes. An API like `swapaxes(arr, 'batch', 'time')` is 10x easier to use than `swapaxes(arr, 0, 5)` -- like, imagine somebody adding an extra dimension in a world where positions are absolutely referenced and where if you get it wrong the program will still run and produce interesting-looking garbage because the definitions of `np.dot` and everything else in the library depend on the shape of the inputs.
A comment would be fine too, especially if it's right next to the type signature, but to do that you'd need to add extra newlines, and the comment would be in roughly the same spot as the type hint, so I don't know that you gain much. Mypy doesn't really like strings used that way, but mypy isn't a great tool anyway, so c'est la vie?
If somebody just wanted to throw that in a docstring I wouldn't complain though. It's definitely more important that the information exist than that it be in a particular place.
If this sounds fun then you can go play with e.g idris or F*.
Runtime behaviour determination: the stdlib [dataclasses](https://docs.python.org/3/library/dataclasses.html#module-da...)
Dataclasses is notable because it's the only example (I'm aware of) of type hints effecting runtime behavior as part of the stdlib.
Compiler instructions: mypyc was (one of?) the first to do this, but Cython actually supports this natively now, and is much more active than mypyc is last I checked.
class Hero(SQLModel, table=True):
id: Optional[int] = Field(default=None, primary_key=True)
And that's fine. I wouldn't necessarily change anything here. Annotated gives you the option of approaching things in a different way, though. class Hero(SQLModel, table=True):
id: PrimaryKey[Optional[int]] = None
I have some use cases where the alternative approach is useful, like quantification of class fields or function arguments.FWIW, `typing.NamedTuple` did this in Python 3.5, three years before dataclasses was introduced in 3.7.
class Foo(typing.NamedTuple):
a: int
b: str
f = Foo(a=1, b="hello")
print(f.b) # "hello" Foo = typing.NamedTuple('Foo', [('a', int), ('b', str)]) >>> @dataclasses.dataclass
... class D:
... x: int
...
>>> D('a')
D(x='a') @dataclass
class A:
a: int = 0
b = 1
>>> A(a=1, b=2)
TypeError: __init__() got an unexpected keyword argument 'b'> To add overloaded implementations to the function, use the register() attribute of the generic function, which can be used as a decorator. For functions annotated with types, the decorator will infer the type of the first argument automatically:
That appears to be the only other case.
Type hints also bring improved completion, which is nice too.
[1] For example, huggingface's transformers library decided to drop support for full type checking because it was unsustainable but decided to keep the types for documentation[2]. There are stubs for pandas, but they're not enough because pandas has a tendency to change return types based on the input, and that breaks quickly.
A mechanism like Haskell's type application seems like it could solve at least most, maybe all of those problems.
The way everybody else seems to be going is strong typing at function interfaces, with automatic inference of as much else as can be done easily. C++ (since "auto"), Go, Rust, etc.
Both mypy and pyright will do that. If your function return type is annotated, they will infer the type of the receiving variable. If you have two branches where a variable can receive two types, pyright will infer the union type. Similar for None.
Example:
a = input()
if a.isdigit():
x = int(a)
else:
x = a
reveal_type(a)
reveal_type(x)
Pyright output, stripped of configuration noise: typetest.py:6:13 - information: Type of "a" is "str"
typetest.py:6:13 - information: Type of "x" is "int | str"
Mypy doesn't allow this. It infers `a` as `int` and rejects the second assignment.The only times I need to annotate local variables are (1) the function isn't typed, so it gets inferred as Any (2) I'm initialising an empty collection, so its type might get inferred as e.g. `list[Unknown]` (pyright; mypy can infer the element type).
Is there something inference-wise that you miss in Python compared to C++ or Go?
PS: The larger problem to me is the inconsistence between pyright and mypy, the leading type-checkers. Sometimes issues are raised between them and they work to achieve agreement, but I believe the two issues highlighted above (unions and collections) are design choices, unlikely to change.
This still makes me seethe. We have pip, poetry, conda, and more. The Python folks knew that multiple incompatible systems would arise from a grammar spec without a behavior spec. And here we are. Python doesn't do anything useful with the types, but third-parties are left to their own devices.
Python typecheckers predate in-language annotations and drove the spec, not vice versa.
[0] https://peps.python.org/pep-3107/
[1] https://github.com/python/mypy/commit/6f0826a9c169c4f05bb893...
[2] https://peps.python.org/pep-0484/
[3] https://github.com/microsoft/pyright/commit/1d91744b1f268fd0...
That said, I don't think this a reason not to use type hints. They're still useful (see TFA), even if they're not perfect.
Mock up something fast, no type hints.
Now, take that POC and make it production ready, by using mypy and pydantic.
(This is also why there is no good HTTP client library in the stdlib, even though the popular `requests` library gets new releases every 6 months with minor fixes only)
Than watch it exploding in production because your "type system" is incomplete and unsound.
In my opinion an unsound static type-system is worse than no static type-system at all. In both cases you need to check everything manually. But without such pseudo type-checking you at least don't get lulled into a false sense of security.
Crystal is compiled with static typing but looks like Ruby. The type specification it uses emulates gradual typing of dynamic languages.
The two languages take such different approaches because their designers have different feelings about static typing. Guido and the Steering Council seem to want Python to be as statically-typed as possible, whereas Matz thinks "static type declaration is redundant" [0].
I remember that it felt rough. I had issues specially with funcions using veriadic types in generics. I also remember having issues with overloading a function: sometimes it would go for the more generic one, instead of going for the more specific one when inferring types.
I managed to solve all of that. Unfortunately, that happened some time ago and I don't remember the specifics, only that it was a fun project to develop. I use it frequently in other projects.
Do not know?
Haskell dies in syntax.
Ex: Programmers value types much more for documentation vs preventing bugs. I had not expected that answer!
That number seems to be from 2013. I imagine it would be much higher today - 10 years ago dynamic typing was at the peak of its hype cycle, and right now it seems to be in the Trough of Disillusionment.
Speaking as a scientist, I'd love to see a round 2 of this work to see how much things are the same vs different. The premise for the work was doing more serious sociological methods could help understand tough phenomena here & suggest new solutions, so bringing in longitudinal analyses would be fascinating.
FWIW, if I remember right:
- Languages like Java, C++, .NET were the most popular. Maybe iOS/Android apps too?
- Buzz around then were Scala + Java (big data), Haskell, Elm, D, and the beginnings of Rust.
- TypeScript was already a year or two in as well. But population-wise, probably still niche.
- That was probably also some of the heaviest Node
Nowadays, we also see heavy rises in dynamic languages:
- Professional data scientists using Python & R. I think Python has become the #1 language for new folks?
- Go looks like a statically typed language... until you compare it to modern C++, D, Rust, etc
So I wouldn't be surprised to see a shift.. but then again, not at all obvious how much, and especially controlling for selection bias..
They never appreciate that type hints make code easier to read and write and navigate and understand and maintain.
Often it's Vim users that have never used a good IDE.
Electon-based IDE seem to require constant access to a power outlet, at least on my laptop.
I've never understood this "battery argument".
I'm not a fan of Electon-based "apps" either, as they're very problematic in all kinds of ways.
But I never could, and still can't, understand the "battery argument".
Battery life is not the biggest issue with Electon (and actually IntelliJ is even a greater power sucker).
The point is: Most people doing software development never "work" anywhere where there's no power outlet!
Most modern trains have power outlets. Likely every coffee place in the western world has power outlets.
If you're not an a safari, or in the forests, or desert there always will be a power outlet near you.
So the whole premise of "I need 10h of battery life" is moot. You actually never need that!
Even if you work sometimes on the go the battery needs only to last as long as you're moving form power outlet to power outlet.
Everything else boils down to: People these days are even too lazy plug in a charger…
Maybe we can invent wireless power to go along with the wireless ethernet?
Personally I would prefer to just use less power. Chips have never been more efficient, and batteries have never been bigger. With the right software, even an older laptop can last all day.
I'd say performance is far from the only reason to consider type hints.
> I don’t know how many people are doing this, but tools like mypyc will use type hints to compile Python code to something faster, like C extensions.
No, the reason to consider type hints is because it makes it easier to understand existing code and write new code that interacts with it correctly. Comprehensibility and correctness are more important than performace.
> Mypyc compiles Python modules to C extensions. It uses standard Python type hints to generate fast code. Mypyc uses mypy to perform type checking and type inference.
> Mypyc can compile anything from one module to an entire codebase. The mypy project has been using mypyc to compile mypy since 2019, giving it a 4x performance boost over regular Python.
I have not experience a 4x boost, rather between 1.5x and 2x. I guess it depends on the code.