I Want a New Duck (2020)
glyph.twistedmatrix.com
glyph.twistedmatrix.com
Admittedly I've never done serious work with mypy (or typescript), so I'm approaching the value proposition of dynamic typing at face value rather than experience. However, it seems like the primary benefit of these languages was ease and flexibility, ableit at the cost of structure. Or said differently, adding mypy feels like trying to get out of a trade-off decision.
This situation reminds me of a talk Bryan Cantrill gave on platform core values, and as examples he gave his interpretation of the platform core values of languages like C, Awk, and Scala: https://youtu.be/2wZ1pCpJUIM?t=349
For me, platform core values that stick out for python would be Approachability, Simplicity, and Velocity. I understand the posited value mypy brings to the table, but it feels in contention with the original core values that made python appealing to begin with.
Oftentimes, the benefits of using Python in a code base that would benefit from static typing outweigh the costs. Especially when tools like MyPy exist, which aren't perfect but help tremendously.
With competing priorities like that, this is the best we can do right now. The other factor is the rest of the team is already familiar with Python. They're mostly infrastructure developers. They're more likely to do everything in Bash if given the choice. I'm not going to teach them Scala or Haskell or Rust in the next three weeks before we need to deliver something. But I can at least teach them Python's optional type hints and mypy.
On the other had we have: https://news.ycombinator.com/item?id=26539508
Sounds like not everything in Scalaland is sunshine and roses.
I think the one connection I'd make back to my original point is that perhaps python becoming the defacto interface for certain libraries is still a reflection of its core values.
This is especially evident with libraries like Tensorflow, which have interfaces for a breadth of languages, and which the core is implemented in C++. The reason people tend to reach for python to call into Tensorflow is still the core platform values of ease of use, rapid prototyping (imo).
Type annotations help a lot. They're not perfect, but for long-running jobs it's a huge help to catch something when PyCharm highlights it instead of 30 minutes into the job.
The reason I'm interested is because I love python and would like to hear differing opinions
For example, consider implementing complex rational numbers. If you have a complex number type, and a rational number type, it should be relatively easy to combine them. Julia does this with literally 11 lines of code in Rational.jl, and 0 lines of code in Complex.jl . In Python, this is extremely difficult, to the point that nobody does it – you would need to restructure the class hierarchy, and you can't do that without access to both.
Another example is the lack of extension methods. For example, if you want to add a bit of functionality to a string or a sqlalchemy engine, you can either add functions (which have a different syntax) or inherit and follow the awful "MyString" pattern.
Python chose to emphasize list comprehensions over map + filter. The problem with that is that you now have syntax that doesn't generalize to other collections, especially custom collections.
Some patterns like async/await, decorators, and with syntax encourage hardcoding decisions when writing a function, as opposed to when using it. This makes your functions less flexible and means you have to write more of them. E.g. consider Julia's do-block syntax[1], which is very similar to Python's with but based on function composition and far more general as a result. Or compare Julia's @spawn to Python's async/await.
[1]: https://docs.julialang.org/en/v1/manual/functions/#Do-Block-...
It's really not difficult; yes, it's more verbose, sure.
> Python chose to emphasize list comprehensions over map + filter. The problem with that is that you now have syntax that doesn't generalize to other collections, especially custom collections.
Python has map/filter/reduce and they work fine. Comprehensions (and genexps) are more concise and cleae for 95% of use cases, I find, so I think the right choice was made there, though I do think people do tend to reach for them sometimes when map, filter, and friends are more appropriate.
> consider Julia’s do-block syntax[1], which is very similar to Python's with
It's more similar to Ruby’s block notation than Python with; it constructs an anonymous function and passes it to another function being called. It is true that with multiline lambdas at all, much less a convenience notation for passing them to other functions, with notation would be superfluous.
> Or compare Julia’s @spawn to Python’s async/await.
Why? They solve different problems. @spawn is equivalent to Python’s concurrent.futures.submit, and it’s buddy fetch() is concurrent.futures.Future.result(). Python has had those longer than it has async/await.
Can you point to a case where it's actually been done? I don't even see how you would do it in Python without modifying the standard library or reimplementing significant chunks of cmath.py or fractions.py . In languages like Julia or Haskell this is possible (and relatively easy) as a third party who's just importing the two types.
> Python has map/filter/reduce and they work fine.
Python's map has a few problems, e.g. crippled lambdas make it less useful and it returns a map object instead of the same type as the original collection. It's better than nothing, but it's noticeably weaker than map in other languages.
> Comprehensions (and genexps) are more concise and cleae for 95% of use cases, I find, so I think the right choice was made there, though I do think people do tend to reach for them sometimes when map, filter, and friends are more appropriate.
One big use case where it doesn't apply is numpy arrays or pytorch tensors. It's very common, at least in my domain, for people to initially write code with lists and then switch to arrays for performance. But despite the fact that they should have mostly the same interface, this requires lots of little syntax changes and it's easy to introduce a bug.
Haskell has this issue too to some extent, linked lists are the default data structure and `map` doesn't work with others, you need `fmap`.
> It's more similar to Ruby’s block notation than Python with; it constructs an anonymous function and passes it to another function being called. It is true that with multiline lambdas at all, much less a convenience notation for passing them to other functions, with notation would be superfluous.
Good point.
> Why? They solve different problems. @spawn is equivalent to Pythons submit from concurrent.futures, and it's buddy fetch() is Future.result(). Python has had that longer than async/await.
I do like concurrent.futures, but most blog posts, documentation etc. recommend using asyncio when either one would work, and the majority of libraries at this point use asyncio. I don't think it's really possible to avoid asyncio at this point.
If I understand the problem correctly, and you are trying to implement Gaussian Rationals, you’d inherit from numbers.Complex, storing the real and imaginary parts as instances of numbers.Rational (you could use the concrete fractions.Fraction, but I don’t think you actually would need to reference the concrete type to do the implementation.)
There’s a bit of boilerplate isinstance-stuff implementing the operations, which is somewhat tedious, but not difficult (I think its less involved than the custom integral class used as an example in the numbers module documentation, because you should be able to get by only handling same-class, Rational, and Complex cases for most ops.
> One big use case where it doesn’t apply is numpy arrays or pytorch tensors.
Good point. It is not one I’ve hit a lot because of the stuff most of python code use is on doesn’t hit those, but I’ve seen that.
> I do like concurrent.futures, but most blog posts, documentation etc. recommend using asyncio when either one would work, and the majority of libraries at this point use asyncio. I don’t think it’s really possible to avoid asyncio at this point.
Static typing is not a job though. It's a language feature. The job is solving technical/business problems.
And Python is a tool, but it's not a tool to do dynamic typing or static typing with, it's a tool to solve problems with.
So, the comment "if you want static typing, you're better off recognizing that python is not the right tool for the job" doesn't quite sit right.
You could instead better say that it's not a good fit for Python.
But why would that be?
Especially since one could easily have said the same for Javacript, but Typescript exists as basically Javascript + types, and people seem to love it.
So there's that.
The extent to which a) is due to implementation decisions made by typescript/mypy vs being due to inherent differences between Javascript and Python is arguable. Certainly there are things that look like unforced design errors in mypy, and the fact that Dart existed (and largely failed) before Typescript shows that it's not just about what language you're based on. But there are also idioms and aspects of the Python object model that seem inherently hard to type nicely, and are sadly too entrenched in the ecosystem to change.
I personally don't agree. I have coded in C# and TS extensively, and while first-class types available in runtime (especially with generics <cough>though not in java</cough>) is super nice, I think the benefits of looseness of TS overweight the costs. Also, you can always go crazy and use zod or io-ts, but in that road there's always the danger of just writing types and not doing any work because "typing is fun"(c).
Note: I'm a python dev, yet even I can agree that the python ecosystem is broken and if not maintained, the language will decline over time.
There is a tendency to think of pragmatism as throwing away core values in order to solve a problem, although generally this is in the use of the word pragmatism in the domains of business or politics and not in the domain of programming language extensibility.
If the one and only thing you want is static typing, sure.
That's, like, never the case in programming, so it's not really meaningful, though.
In real world programming scenarios, you are always dealing with balancing multiple dimensional problems and “python, with some degree of static typing” is a reasonable component of the solution for lots of them.
> Or said differently, adding mypy feels like trying to get out of a trade-off decision.
No, considering python’s available static type checking (whether mypy, pyright, or whatever, or even a combination) is turning a big coarse trade-off decision into one with finer-grained options, not avoiding it.
> For me, platform core values that stick out for python would be Approachability, Simplicity, and Velocity. I understand the posited value mypy brings to the table, but it feels in contention with the original core values that made python appealing to begin with.
IME using Python’s typing, usually incrementally added, it's not particularly. In fact, it's an enhancer to velocity as the code base evolves.
Here's the song for the curious https://m.youtube.com/watch?v=eOL2q8leiLw
From 1985.
Edit: Answered my own question via Google: https://consequenceofsound.net/2019/06/weird-al-yankovic-mic...
Random: as someone who makes heavy use of both Python and Typescript, it pleases me to no end to see a great parser and ast for python implemented in typescript.)
My personal website is being rebuilt on Deno at the moment! Writing type level operators is a lot of fun haha
Both Go and TypeScript were right there showing you how effective it was.
For the record, duck typing is the way. Mypy is a curiosity to me; I'm happy those who need it are getting that itch scratched; I'm happier it's boxed away from what I work with.
I don't think this has to be the case. If it got serious investment and the Python community was more open to making ergonomic syntax improvements so type annotations could be natural, we could see improvement. Similarly, I think it would need to figure out a saner solution to finding/loading type annotations, because the current set up and error messaging are immensely painful. I also don't know if Mypy will ever be sufficiently performant so long as it's written in Python (I think people really underestimate the difference between instant feedback and a delay of several seconds). Having good editor support would also drive adoption.
If Mypy were able to improve on all of these distinct problem areas, I think more people would opt into type annotations naturally, but it certainly feels like Mypy is being treated as a thing to pacify the people who whine about static typing (of course I don't think that's the real intention, only how it comes across).
Inside a function you need modify, you are passed a foo. What can you do with it? Take the length? Add a number to it? You don't know unless you know the type, and then the only place that info exists is in your head. Why not write it down in a way that can be checked and therefore never get out of date, like comments or some weird reincarnation of Hungarian notation?
As a bonus: Editors will tell you when you are passing incorrect types to functions before you even run your tests. No more strings accidentally treated as lists!
I am now absolutely convinced that types are important, even in python.
If mypy is doomed to remain marginal it is because it picked a poor type system for the language it tries to support, not because stricter typing generally is the wrong choice for that language.
One difference may be if you primarily work in your own projects or work with other people’s projects. And how skilled the other developers are. In my own projects I instinctively know the types, but not in other projects, and having the types documented allows me to make changes much more quickly.
Type-advocates throw this around all the time as if it's not the most patronizing thoughtpattern. Yes, I work on a good sized project with another 15 engineers using Python for the backend and angular for the frontend. If scale were going to reveal something dramatically different than whatever toy algorithms you might imagine I'm playing with...it would've happened by now.
What I meant was more along the lines of “do you regularly jump into the deep end on new projects that were started/maintained by someone else?”
Doing that is what made me a believer. Jumping into my current project, taking over from some else, was a nightmare without types. I had no idea if something was a string, a dict, a dataclass, a custom class, etc, with only vague hints based on how it was being used. These obviously had “types”, and the functions were designed for only one set of types. I just couldn’t remember what.
“Your own” doesn’t necessarily mean small or toy. But more about consistently working in familiar code bases where you are already familiar with what types are flowing around.
I think of these kinds of types as enforced documentation. It’s not for you (necessarily), it’s for other people.
The overhead is as much or as little as you like.
What's more, with static typing your IDE can provide you with more helpful guidance by immediately showing you the valid method names associated with an object, type-checking their parameters as you use them, and suggesting possible parameters of the required types from the variables currently in scope.
With dynamic/duck typing, you give all that up -- for no clear advantage.
Use static typing in a project of significant size. Always.
And introducing complexity you have to resolve before your program even runs too! Complexity you have to solve. Static typing systems require you to do more design upfront. And more importantly, in your head! The most you can do is run your compiler and it'll tell you if you got it right. If not, just try again!
Dynamic systems will encourage you to code flexibility from the beginning. You may have to run your system to expose a class bugs...but so? I was going to run it anyway.
All these conversations are just about moving work around and when it gets done. Most problems are solved equally by a dynamic or static type system. However, static type systems inherently require more boilerplate. At the end of the day, you just rarely need it to do anything of importance.
The formalization step also allows you to encode business logic, where it would not otherwise exist. Suppose you have a system with documents. Every document has a length property...today. So in an untyped codebase, you've been doing just fine assuming it exists, and "just running the code" has found no issues. But technically, it doesn't always exist, and this fact isn't encoded anywhere. It breaks. If you had just defined a document type, then you could simply update the length type and let mypy tell you literally all the places where the bug would arise. Or, it would have allowed you to avoid making that assumption in the first place, rather than trying to make sure every developer memorizes this useless fact and writing tests every time the length property is touched.
You've baked your conclusion into your argument. But by this logic, sure, everything is typed. I'll buy it. A static type system just makes you explicitly type it out for every single thing and again for every single thing that might operate on that first thing. It's pedantic, filled with boilerplate, and plenty of promises about problems the system solved that were just invisible before. Hmm...those might be the marks of snake oil, now aren't they?
Consider a iron foundry. Are there type checks on the iron? No, the entire process is duck-typing. Why? Because that's how reality works. You can specially shape your iron ingots to key into your foundry. What's that? Duck typing. The firing profile will follow certain characteristics. Did it get those characteristics from the ingot? No, it duck types it and starts firing it as if its iron. If it isn't iron, the process will throw an exception. And that part has all sorts of safety checks around it. The most you can do would be some type of chemical check on the ingot before starting. Almost like a typecheck before proceeding in a method.
Aristotelian vs. Platonic approaches. The problem with the Platonic approach is that it requires you to be an oracle. Static typing is like trying to encode logic at the molecular level. This atom WILL NOT bond this atom. I mean, sure, it can work...but there's a lot of atoms and a lot of interactions. Don't define what you don't need.
Yes there are, it's called physics haha. The iron foundry analogy is not great imo. Programming is more akin to building an engine, but also building the tools that are used to make that engine. Because of the physical nature of these tools, you get a simple static type system by default: you cannot put a screw into a hole not meant for the screw with that thread. Now imagine you didn't have this guarantee. Any tool could technically be put anywhere, and you wouldn't know you had it wrong until you ran it and it killed you. Or even worse, it seems to run fine until you go over a bump, then it kills you. That's duck typing without a type system.
And god forbid you need to add more engineers to your engine project. Nothing is documented, and they have to infer the properties of the tools based on how they are used today.
That's half of duck-typing. You're missing the exception response where safety is enforced. A static type system fails with the wrong type. A duck-typed system responds with something as sensible as it can reason with what it was handed. Must often, that's a hard failure because most code is brittle and doesn't need the robustness (i.e. don't build what you aren't going to need).
That's before we even get into the philosophical issue of, who are you, the lowly library implementor, to say those aren't something appropriate to abstract over in my application?
Scientist.quack really shouldn't be returning plain strings, and neither should ducky.quack if the data in them is contextually different.
Things should only become basic types at the last possible moment.
This is one of the reasons it is so rare to actually have confusable types in a real program.
The interfaces matched, the compiler accepted it, but the application blew up at runtime.
Extending this, if the class has a `partial_fit` method, sometimes I want to use that, and fallback to `fit` when it doesn't.
And the official one obviously couldn't be used.
$ cat test.py
from dataclasses import dataclass
from typing import Protocol
class Ducky(Protocol):
def quack(self) -> None:
"Quack."
class FringeScientist:
def quack(self) -> bool:
return True
@dataclass
class Duck:
quiet: bool = False
def quack(self) -> None:
print("Quack." if self.quiet else "QUACK!")
def duck_war(aggressor: Ducky, defender: Ducky) -> None:
aggressor.quack()
defender.quack()
print("The only winning move is not to play.")
duck = Duck()
asshat = FringeScientist()
duck_war(duck, asshat)
$ mypy test.py
test.py:32: error: Argument 2 to "duck_war" has incompatible type "FringeScientist"; expected "Ducky" [arg-type]
test.py:32: note: Following member(s) of "FringeScientist" have conflicts:
test.py:32: note: Expected:
test.py:32: note: def quack(self) -> None
test.py:32: note: Got:
test.py:32: note: def quack(self) -> boolLooks like mypy has discovered interfaces. Lovely.
A Protocol is just the same structural subtyping (duck typing) Python has always preferred.
I use static Python types a lot at the application layer, where things have no need to be particularly extensible and having strict definitions for how things are put together brings clarity. But at lower levels of the system, Python's duck typing really shines because I can solve ten slight variations on the same problem with the same code.
The same could be said of Typescript, or Clojure+spec, or really any other language where dynamic types are given to you for free (unlike, say, reflection in Java) but types can be added wherever you find them useful.
Having had this experience now it’s much easier to appreciate the practical benefits of systems like Rust and Haskell, where the type checker is doing real work.
Getting that cheaply while still having types unchecked is a nice compromise between improved functionality and not changing the language majorly.
* To the extent that it enables some static analysis on python code, it has often caught a number of errors I tend to make while writing code
* It serves as a light form of documentation that tends to stay more fresh because there's an automated way to check its consistency (run the type checker)
I don't expect that it will do as perfect a job as Rust or C++'s type system or even that it would be there for the same reasons. On the other hand, it's nice to just have more info about what the code is expected to do in the form of type annotations.
Mypy started out as its own language with a Typescript-like relation to Python, evolved to being fully python compatible (using typing comments) after consultation with Guido (then-BDFL) and then python added annotations to support projects like mypy. It wasn’t just the that annotations were created and then mypy happened along to take advantage of them, supporting mypy specifically (without only supporting mypy) was a key motivation for annotations.
The only time I've ever used `Protocol` is to define a type that makes it explicit that I need an object to have a `__str__` implementation:
@runtime_checkable
class Stringable(Protocol):
""" Protocol type that supports the __str__ method """
@abstractmethod
def __str__(self) -> str:
""" Return the string representation of an object """
I've since learned that this is redundant, because all Python objects have an implicit `__str__` if one isn't specified (IIRC). When I didn't know this (when I implemented the above), I didn't use an ABC because obviously I can't guarantee all objects (e.g., those outside of my control) are subclasses of said ABC. The cases where this is true are vanishingly small, especially when you go big on dependency inversion.Abstract base classes require everything to extend from a base-level object, and also inherit it's default implementations. This is a nominal (name-based) mechanism.
Protocol (structural-based) subtyping instead only declares what methods something has to have, without requiring it tie itself to either to a concrete parent class and its behaviours.
Edit: since you partially address this, you ask why you'd use structural typing rather than an abstract base class. I'll turn it around and ask why you would use an abstract base class, which requires modifying the children accordingly, when you could instead use a structural type. The usual preference comes down to whether you want to be explicit about implementing it or not, and whether you want to pull in default behaviour.
Thanks :) I definitely prefer things to be explicit; regardless of it being a "Zen of Python" mantra. That said, I do see the value of `Protocol` for non-annotated/legacy code, in that it gives you a convenient mechanism for this. I would be worried about it becoming a crutch, if overused, but definitely better than no annotations at all!
(I don't understand your point regarding an ABC's default implementation. Isn't the point of ABCs that they're not base classes, but interfaces that define the methods required without any implementation?)
The very definition of an abstract base class precludes the existence of (meaningful) implementations. If the base class has implementations for the methods expected to be overridden (beyond no-ops or throwing some exception around the lack of an implementation), then it is definitionally not abstract.
That is: a protocol and an abstract class achieve the same thing: they declare the existence of methods, and defer implementation of those methods to implementations / child classes (respectively).
Interfaces are great, but if you have multiple inheritance as your hammer (as you do in Python), then it's perfectly reasonable for abstract class inheritance to be your nail.
If the not-overriden method isn't commonly called, it's possible for the non-meaningful / raise an exception version to make it to production. Not so if the error happens at typecheck time.
No, it doesn't. It precludes meaningful implementation of the complete interface, of course, but it doesn't preclude meaningful implementation for some methods, which depend on the other methods do which meaningful implementations are not provided.
The fundamental issue is that Python doesn't really have a way to enforce this, so there's ultimately no such thing as an abstract class in Python - there are only classes that act like they're abstract (by replacing all implementations of meant-to-be-overridden methods with no-ops or thrown exceptions).
When (as to make this even worth discussing, one must) one includes the typecheckers (which are technically external) in “Python”, that's not true; both mypy and pyright enforce that abstract classes (either those explicitly declared as abcs, or derived classes with any non-overridden declared-as-abstract methods) cannot be instantiated.
The general OOP concept of an “abstract base class” might, the particular Python implementation does not, because, as well as classic inheritance, it supports the concept of “virtual subclasses” that are registered with, but do not actually inherit from (and thus do not use the implementation of), an abc.
It shares with structural subtyping that it only defines a mandated interface, but with nominal subtyping that it requires an explicit statement of intent to conform to the interface before an object is considered to conform.
Except that most typed languages are far faster than python. Retrofitting existing python projects with types makes sense for reliability, but starting a new project with python and types makes less sense, imo.
People use Python for its ease and speed of development, and extensive ecosystem. Things that need to be fast can have the appropriate effort expended to execute as fast as necessary.
Of these, I think only the ecosystem stands out anymore. For work purposes I wouldn't be able to move away from Python anytime soon, but if I could pick any language for new code it wouldn't be Python. If I had to pick the closest competitor, Julia seems like a great Python-like language in terms of speed and ease of development with much better typing and performance, but with a less extensive ecosystem. Nim might also fit that bill. If ecosystem was important, Go probably fills a similar space without being quite as steep of a learning curve as some of the other languages.
Well as a Perl fan I certainly don't disagree with that conclusion ;)
But yeah, you're definitely right that there's immense value to being able to quickly prototype.
That said, if you care about the performance improvements that typing can give you with Mypy, you might want to look here:
https://github.com/mypyc/mypyc-benchmark-results/blob/master...
It won't be going toe to toe with Rust any time soon, but a 4x to 17x speedup is nothing to sneeze at.
Most of these things can be improved upon, but progress feels slow and there are so many other serious issues in the Python ecosystem (performance, package management, etc) that I've given up (after 15 years of daily use). It makes me sad, but Go and Rust seem to offer better tradeoffs for the things that I care about.
Like, we all know about the expression problem and how interfaces are hard to modify after the fact...
But are test mocks and stubs not literally the prime use case for interfaces at the syntax level?
No. e.g. Java interfaces are just as nominal as classes.
> But are test mocks and stubs not literally the prime use case for interfaces at the syntax level?
Also no. If your language has structural interfaces it is often preferred to inheritance as a tool for polymorphism. Even in C++ where templates are not the easiest tool to wield it's often still preferable to inheritance for many reasons.
This is very literally how people who write in languages without structural types do test stubs (and often dependency injection, and other mechanisms that need to select from among several implementations of something).
Having to fight the type system just for testing seems silly. Just have every class's public bits define an interface and the class's name in code always refer to the interface.
Then testers can just implement ClassName with no fanfare and you stop having to write "test-friendly" classes.
If memory serves, Twisted hasn't been too bad. SqlAlchemy is pretty well a nightmare though. I think I saw that the latest release has type stubs. That will be nice. PyQt is also less than ideal...
Python’s main implementation has a bytecode compiler (that’s where pyc/pyo files come from.)
Mypyc compiles type-annotated python code to native-code python extensions.
Cython is a native-code compiler for a superset of python.
Both Mypyc and Cython are used for widely-used code.