Python 3 Types in the Wild
neverworkintheory.org
neverworkintheory.org
The paper referenced is a year and a half old and type checkers have fixed/added a lot since then. Also, I'd guess that more projects use Mypy than Pytype and fix errors reported, so that may be another reason more pass Mypy even without annotations (though conservatism does play).
this type is impossible in Python
I could maybe see something like this as an intermediary result of parsing some horrible, nested obscure message format (except the dict, that baffles me unless the key-value types are reversed?)
If the need exists for a list of lists then it exists for a set of sets.
An example of this is finding all combinations possible non ordered groups of letters given a collection of letters: abadfjiik.
A typical algorithm will produce dupes. But if your accumulator is a set of frozen sets you can produce a result with no dupes.
FirstNameType = str
(But don't use ints for postal codes, even if the postal codes are numeric)Maybe in the US. In Canada, they are 6 characters of letter-digit-letter-digit-letter-digit (in regex: /((?:[A-Z][0-9]){3})/)
FirstName = NewType('FirstName', str)
PostalCode = NewType('PostalCode', str)
[1] https://docs.python.org/3/library/typing.html#newtype[== do we have something like Java???]
I use type annotations solely as means of documenting code, with the added benefit of having autosuggest work all of the time. I am one of those guys who put a lot of comments in their code.
You can have a dictionary which maps to strings, you can have a list containing sets of frozensets of ints, but you can’t have a dictionary of unhashable types.
Python 3.8.5 (default, Sep 4 2020, 02:22:02)
[Clang 10.0.0 ] :: Anaconda, Inc. on darwin
>>> d=dict()
>>> l=[1,2,3]
>>> d[l]=42
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: unhashable type: 'list'Apparently there are ways to monkey patch core types, so you could probably add a __hash__ method to list…
That's often a mistake even in typed languages like typescript. Satisfying a static type analyzer isn't enough. Real data is whatever it is going to be. If your types are meaningless at runtime, you can easily encounter type mismatches.
Static typing as implemented in typescript and python type hints is a tool for engineers instead of a tool for systems.
foo=PydanticFoo(**request.json())
Will try and validate a json request against the implied schema provided by the type hints on PydanticFoo while constructing it and will throw if it fails.Pydantic is amazing!
What happened is I started losing out on the advantages that Python provides as a dynamically-typed language without truly gaining those of statically-typed ones. Then I realized, type hints are just that: hints. They're a form of documentation. In some cases, they're very helpful, but in other cases, it doesn't really make sense to use them. I don't need to appease mypy. The type hints are their for the benefit of myself and other developers.
Probably just used to writing poor quality (immediately unintelligible to someone other than the author but functional) code then.
(Similar to when you something you wrote is hard to test, it’s probably just poorly abstracted code.)
One example is that variable re-assignment can cause type errors in mypy but re-assigning a variable to a value of a different type is an extremely common and reasonable thing to do in Python. Another situation where strict typing can be borderline hopeless is when parsing e.g. very dynamic content like json. Sure you could encode every possible situation using types but this would basically defeat the purpose of using Python for such a task.
But it's a wart worth the conciseness, at least, until a better idea comes along.
Maybe if type hints could be optionally enforced as a language feature, I'd feel it was more integrated.
> as evidenced by only 15% of repos passing mypy
That is no evidence. Of course, if you don't have mypy in CI (or pre-commit hooks or whatever), your repo will not pass mypy. But if you have it in CI, then you can rely on them being correct.
> tools like pytype or mypy will never capture all the complicated hacks possible
But they can either enforce everything being correctly typed (if you go to really strict settings), or yes, you will have some code untyped, so you will be careful around that part. It's not like all or nothing.
In practice, most modern statically typed languages maintain soundness through a combination of compile time and runtime checking.
Rust, C, C++ - all do not keep around this information.
You're right that the JVM languages (and thus, naturally, C#) do it differently. I think it is a bit silly to name out each of the JVM languages separately.
I think basically every managed statically typed language ends up carrying around runtime representation and uses some amount of runtime checks.
AFAIK - and this might be either dated or just wrong - the two places GHC doesn't fully erase are 1) polymorphic recursion (where if we squint it's probably not unreasonable to treat the dictionaries as type tags), and 2) when the programmer explicitly asks for run-time type information with Typeable.
However, such matches are considered valid code: "non-exhaustive match" is a warning by default, even if most people advise to turn that warning into an error.
In other words, the `Match_failure` exception is part of OCaml non-typed semantics and there is no soundness issue involved here.
For instance
let f x = match x with
| [] -> ()
is translated into (is_empty/267 =
(function x/269 : int
(if x/269
(raise (makeblock 0 (global Match_failure/18!) [0: "r.ml" 1 17]))
1)))
where you can see the exception being build and raised in the `then` branch of the test.
Contrarily, the total function let is_empty x = match x with
| [] -> true
| _ :: _ -> false
becomes (let (is_empty/267 = (function x/269 : int (if x/269 0 1)))
And since the function doesn't need to handle the failure case, it doesn't have that `(raise ...)` case.Another important point is that type system information is only needed to check the exhautiveness of pattern matching in presence of GADTs. Otherwise the exhautiveness of pattern matching can be checked using syntatic criteria on the pattern matching and type definitions.
In the GADTs case, the type system is only used to remove the failure branch. For instance,
type 'a t = A: int t | B: float t
let always_a (x:int t) = match x with
| A -> ()
| _ -> .
is translated to (let (always_a/270 = (function x/272[int] : int 0))
because the typechecker can prove that the `B` case is impossible.
In other words, GADTs are yet another instance where the type system can be used to eliminate dead code in the untyped IR.Makes sense to me.
I mean, machines serve people (for now!). So in addition to type documentation, type annotations are useful for things like "jump to definition" and "find references" and automatic refactoring. Moreover, the utility of annotations for type documentation depends on the correctness of those type annotations--if you don't have something like Mypy, then those inevitably annotations grow outdated over time and incorrect type annotations are even worse than none at all.
> tools like pytype or mypy will never capture all the complicated hacks possible in the Python type system
Technically you can represent any complicated hack if only via `any`, but that's just pedantry and we all know what you mean. Even still, it's not like "well you can't represent everything, so there's no point in trying to appease mypy" (which may not be your intended implication--I can't tell). The benefits of appeasing mypy are proportional to the amount of your code that is properly annotated--this is the thesis of gradual typing.
> as evidenced by only 15% of repos passing mypy
This very likely indicates that 85% of repos aren't running mypy regularly, or they're in some transition state. Certainly more than 15% of real-world Python can be described via Python type annotations.
Compare it to Typescript which is a similar sort of tacked on type hinting system - Typescript is basically always machine checked.
It was simple, elegant, perfectly cromulent Python code. However, there was no way to specify the type of the dicts because they were so heterogeneous.
- - - -
I have to say, I was a Python fan for ages, used it professionally for over a decade, so this isn't an outsider's opinion: Python is kind of over. It won't die any time soon, but it has been "improved" out of its sweet spot of applicability, and no longer has a compelling story for adoption going forward. Every niche Python is known for serving well now has even better options:
- Rapid Application Development: Nim, Red
- Readability: Go, Nim, D, Zig
- Science: Julia
- Large multi-dev projects: Java, Go
- Systems programming (python glue): D, Rust, Zig
- Web app backends: Erlang/Elixir
And so on...
Really, as much as I still love Python (and I do) the only things I feel it's appropriate for these days are small scripts that need to remain flexible (in other words, that are likely to be modified often), things that are too complex to be a CLI command, but less complex or more dynamic than, say, an email client or RSS reader. If the type domain of the program doesn't change rapidly or unpredictably then you don't need Python's dynamic typing. If you need static typing I think you should use a language that supports that out-of-the-box, not tacked on later.
Yes I am quite sure for any specific niche there is a language that is better suited than Python (though I would argue with a lot of your choices there). But let's say I want a readable language that is good at data science, web backend, and provides a system glue. Python handles that much better than all the others you mentioned.
Also the one choice I am going to explicitly call out is Rust as a good systems glue language. In fact a very common practise I see now is all the good Rust libraries are getting Python bindings and made available on PyPI so all that power of Rust for systems programming can be used in Python as the glue language.
A small startup might do well to make their MVP in Python, but as the code grows the implicit costs (of using Python) do too.
- - - -
In re: Rust, sorry I wasn't clear above. I don't mean that Rust is a glue language, I mean that people write e.g. grep replacements in it and things like that. Python does systems programming by being glue, Rust does it by being, well, Rust. It makes sense to me that Rust libs would get Python wrappers, but it also seems to me that that adds to my argument: Python is good for small glue, but crunchy things (like grep) should be written in e.g. Rust or Go or something.
(I feel I should mention in this connexion my favorite bit of obscure Python trivia: Python was originally meant to be the shell language of the Amoeba distributed OS! https://en.wikipedia.org/wiki/Amoeba_(operating_system) )
- - - -
One other thing about Python is that the packaging & distribution "story" is ridiculous now. The people in charge of that call themselves the Python Packaging Authority (which name, given what they're doing, reminds me of Brazil the movie) and they seem to me to be running amok, cargo-culting the crap out of what should be a pretty simple and straightforward problem. I could go on but I feel a rant brewing, so I'll cut it off there.
It's not just the PyPA folks that are having problems packaging and distributing Python. The Conda folks ship Tkinter in a broken state for five years now: https://github.com/ContinuumIO/anaconda-issues/issues/6833 That's the default GUI system that's in the Python Standard Library.
Compare and contrast with Rust's Cargo, or Nim's Nimble, or Erlang's Rebar, etc.
You can also export Nim impls to Python easily, though I very much agree Python's time in the sun is over (and I say this having been a user since ~1993). It now seems more a generator of problems/complexity than of simplifying solutions. Nim can also compile to Javascript letting you do both front- & back-ends in Nim.
Science is an almost unbounded fractal set of human investigations. Nim is usable for science (e.g. ask questions here [3]), but the ecosystem is undeniably much smaller than Julia or Python's. (Julia's is much smaller than Python's, and for some things R can be much bigger, actually.) So, a "depends what you need" caveat just has a much stronger "caveat force" here.
What "large, multi-dev" needs is honestly harder to say/more subjective/maybe manager-subjective. If it is "cheap, just out of school programmers with which to replace expensive, departing mess-makers", I think Java may have always had more plentiful options than Python. ;-) This is closer to a human-organization-complete kind of concern and almost intrinsically "very dependent upon circumstances/context". I just didn't want to ignore it. At least Stefan Salewski [4] thinks Nim is suitable or teaching people introductory programming, though.
[2] https://github.com/c-blake/cligen/blob/master/cligen/dents.n...
I might still write a prototype or proof-of-concept web app in Python, but for anything more serious it seems to me that the advantages of Erlang loom large (and that's before you get to Phoenix LiveView: https://www.phoenixframework.org/ )
Silly, Python was always second-best at everything and promoted as a prototyping language from the beginning. But that's more powerful than you've given credit for. You get extreme productivity up front, and a growing number of exit strategies for the fraction of projects that outgrow it.
> a growing number of exit strategies for the fraction of projects that outgrow it.
I like that formulation. Cheers!
[note: i am a java developper :)]
Its because it's the language of choice for bad programmers who dont learn languages with any amount of depth. It's getting associated with mediocrity because it's so easy to pick up and low quality coders turn to it.
Your job solidifies this point. Literally, they need to hire someone to type annotate their code base because they don't know how? How low can you go?
Hope you can fix it. However, be prepared for other kinds of issues, which arise when there was not much discipline in writing the code.
There are few conventions that are consistent across domains. As a concrete example I worked on a code base where functions from some libraries wanted angles as radians and the others wanted them as degrees. Both where internally consistent and followed all the relevant conventions for their domain, however when they intersected it was a pain.
That’s my current plan of study. But as I said earlier, I have some background in static typing in Java, and I might find the current state of the art in Python not to be adapted to my idea.
[1]: https://en.m.wikipedia.org/wiki/Functional_Mock-up_Interface
php?
Sorry, couldn't resist.
On the more serious side, what you're describing about python I feel happened to PHP long ago. I wonder if Ruby is next. Hope not. I'm doing mostly Ruby these days.
Sure, and Java went through this phase too, and I'm sure so have plenty of other languages. It's the stage at which a language has reached mass adoption and it is inevitable, but it doesn't mean there's anything wrong with the language.
Neatness is in the eye of the beholder. Type safety is a quantitative metric of improvement. A program with type checking is more safe then one without.
A quantitative improvement cannot be argued against. A qualitative one... Well, personally I find types to be neater than no types.
How many type errors in a language that is not type checked vs. how many errors in a program that is?
The numerical difference is quantitative. You can therefore do data analysis on this. Though a type safe program should in theory have ZERO type errors during runtime.
Putting it another way -- a type safe program, in the limit, should have an infinite number of type errors. :)
Dynamic typing is not a deficiency, it's a design choice. SICP for example, is filled with small examples of things you can do with dynamic languages you cannot trivially do with static one. There are certain styles of solving problems that suit dynamic languages particularly well, and anyone programming in a dynamic language should understand and embrace these.
There are of course drawbacks to dynamic typing, as is typically the case you exchange predictability for flexibility.
If, for your given problem or preferred style of programming, you want types there are so many fantastic languages around today that have much better type systems than what was available when Python was first gaining popularity.
Programming well in a statically typed language is fundamentally different than programming well in a dynamically typed language. Typing in python will never be powerful enough to allow you to "think in types", and if it were then python really wouldn't be python anymore.
I’m curious to see how different an equivalent solution might look in a statically typed language and if any features like type classes or row polymorphism close that gap.
If you pick a line at random and try to understand what it does or why it's failing, you will have a hard time. Python (by design) is so mega ultra dynamic that each line could conceivably turn the world upside-down. Most code I've seen doesn't take advantage of that dynamism. But still - it's tricky just reasoning about stuff like: what kind of values that call could return?" or "what kind of an operation the index really does here?"
I worked in what I would imagine was a pretty typical mature medium-size python codebase.
Typically, it was never too hard to figure out what a type was. Sometimes types were passed through several layers of functions and had the same short variable name through all of them - quite a few times I had to click around quite a bit to figure out what "f" was. Especially when I could infer it was some kind of file-ish thing but whether it was a file object, a string, or some composite type that held one of the two as a member, or what?! was sometimes difficult to determine, especially if the actual use of the parameter wasn't immediately obvious either.
When we updated to python 3, I started leaving type hints both in old code as I answered these questions for myself and new code as I created what otherwise would've been new questions for my future self. And I noticed two things happened:
* annoyance went down as I fixed the most well-trod of these problems
* I became less averse to using slightly more complicated types
I think python pushes you to keep things really simple and not use custom types unless really necessary. This is, overall, good? I do think Java developers, for instance (of which I consider myself one), generally reach to create datatypes that are basically just a simpler wrapper or a 2-tuple of collections or other primitives, and it can make the code a bit annoying to grok.
But the trouble is, in un-type-hinted python, I already start getting nervous about things like: [('2022-03-18', "something"), ('2022-03-19', "something else")]. And if your data content doesn't make it obvious (or at least somewhat guessable) what it is, it can make it hard to grok in a slightly worse way than the Java code would be.
In python 2 I'd usually make a namedtuple in these situations, but oftentimes I felt a bit weird there because I'd usually reach for it in lightweight situations when I feared they were becoming more complex.
But finally, in python 3, I feel like I'm generally happy with, in this order:
1. just use plain primitive types. no type hints. we all know what's in dates = ['2022-03-18'].
2. just use a type hint. I feel better about a Tuple[str, Dict[int, int]] if it's type hinted than not.
3. Use a namedtuple. This puts names onto the fields. so maybe my Tuple[str, Dict[int, int]] becomes a MyEntry(token: str, settings: Dict[int, int]) or something.
4. use a fully-fledged custom data type class.
I come from a massive hard-realtime system codebase that's mostly in Python and there's a lot of moving parts and complex data types. Type hints, `typing.Protocol`, and `dataclass` are all godsends for having any sort of sane, human parseable structure to the code. Being able to navigate to type definitions with `gd`/ctrl + click is massively helpful.
Not surprising. In my experience, upgrading or downgrading Mypy always breaks type-checking on real code bases: older versions because they fail to infer types or have bugs, and newer versions because they are stricter or error on '# type: ignore' that are now useless (but were needed to work around bugs that are now fixed).
And there are 10 Mypy version released over the last 12 months, so projects probably used different versions than the one used by the authors of the paper.
While the article does not mention when they downloaded the repositories, I am guessing they started in August 2019, given section 3.1 (with no specified end date, besides the paper's publication in November 2020). And the single Mypy version they tried (0.770) was released in March 2020.
Python is not Perl, it's very strongly typed, has operator overloading and close to no "special syntax" (i.e. Perl's sigils, etc), so it's kinda hard to follow code sometimes. `typing` helps a lot.