Statically-typed error handling in Python using Mypy
beepb00p.xyz
beepb00p.xyz
* Third party libraries are still not typed, which sucks
* Inference is weak. Sometimes an 'if' statement narrows accurately, sometimes it doesn't.
* Generics are extremely confusing, moreso than any other typed language I have used. Any Generics are even more confusing.
* Fundamentally, Python is structured such that the concept of interfaces are coupled tightly to implementation of classes, which is severely limiting. As an example, a module A can declare a class, a module B can declare an interface (Protocol), but module B can not implement that interface for A (without modifying A's code).
* Errors from mypy are still very very light. There are very few "Try XYZ to fix it", and it's just a single line of read text that amounts to an assertion.
* Still weak support for JSON (use Dict[str, Any])
* Slowly improving but still not amazing support for Self types
I could go on.
Mypy is awesome but it still feels very immature, and as it should since it's pre 1.0. It does not come close to matching the experience of other languages with static types.
* Regarding errors: there are few quite recent flags: ' --pretty', '-show-error-context' and '--show-error-codes', they make it a bit more pleasant. You can see me using them in this section: https://beepb00p.xyz/mypy-error-handling.html#container
Thank you for letting me know about those flags, I had not heard of them before!
Python 3.8 added TypedDict which finally helps this particular shortcoming.
It's nice to see strong progress in such important areas. My post is probably overly critical, but I write Python almost every day, and maintaining a highly-typed library can be demoralizing at times.
from mypy_extensions import TypedDict
https://github.com/python/mypy_extensionsOptional keys are also a nuisance.
if "optional_key" in my_dict:
do_something_with(my_dict["optional_key"])
Or: optional_value = my_dict.get("optional_key")
if optional_value:
do_something_with(optional value)
Works as expected. Nested optional keys can be a bit annoying, although: my_dict.get("optional_parent", {}).get("optional_child")
seems to work.If you have a dict with some keys which are optional, you need to create a separate subclass (with `total=False`) just for those optional keys.
With TypeScript, I can just use `key?: type`.
https://mypy.readthedocs.io/en/latest/more_types.html#mixing...
While mypy and type hinting support in PyCharm and VSCode is great, it's not the seamless experience as with typed languages and the lack of (optional) runtime typing still allows all sorts of bad things to happen. A "safe" mode for Python, where typing and perhaps some other stuff is checked at run-time, might be an interesting advancement for those in the community building larger or mission-critical systems.
If you write str1 == str2, it compiles but probably does not do what you want (object identity vs string comparison).
(edit: to clarify, this does not mean Java is not a good language overall, every language has its own set of shortcomings and as someone points out in the replies, this behavior is consistent with everything else in the language)
Frankly, if this is the most meaningful complaint you can come up with, it might be indicative that you haven’t learned enough about it to have an informed opinion.
In fact, from a certain perspective the two flavors get treated identically. Comparing two integers with ‘==‘ will be true if the contents of the variables are the same, and comparing two references will be true if the pointers they contain are the same value. It just so happens that for primitives identity and logical equality are the same thing.
The problem is that in Java, there are no user-defined value types, only primitives; and primitives can't have methods. So if it were a primitive, you'd have to write "String.length(s)" etc. Also, all other Java primitives are basically bit sequences that are interpreted in one way or another, but that wouldn't be the case for strings.
.NET/C#, though, doesn't get that excuse. It totally could have defined String as a struct with an internal char[] field, and then there would be no question of value/reference equality for its ==. But that would also mean that you couldn't use null for strings - and they didn't have nullable value types back then. I suspect that, plus the overall mindset carried over from Java, is what won the day.
Side note: there's no hard and explicit distinction between reference and primitive types in C itself, nor in C++. If Python is "everything is an object reference", and Java is "everything is an object reference except for primitives that are values", then C++ is "everything is a value, including object references". Thus, there's no ambiguity with == in C++ - it always compares values, it's just that sometimes those values are pointers.
I think sometimes that perhaps the implicitness of object references that is so common to OO languages, that I think was introduced by Smalltalk, is a mistake. It's interesting that Simula-67 didn't have it - although it was very Java-like otherwise, having only primitive values and references to objects (i.e. no objects as values, like C++), it distinguished the two consistently by using distinct equality operators (= vs ==), and even distinct assignments (:= vs :-).
Or, alternatively, treat everything as an object reference, but make == do implicit dereferencing as well, as Python and VB do; and provide a completely distinct operator for reference equality, such as "is". Python has a problem in that regard in that it has a default implementation of == for all classes that compares references, and so it ends up used as reference equality in practice sometimes. It would be better to have no default for == at all, just as there's no default for other comparison operators; this is what VB does.
It might also be best to stop talking about value and reference types altogether - what's actually important is the presence or absence of identity. Then everything can be an object reference, but not all object references can be compared for equality (and in practice, under the hood, the implementation can just skip the references and copy the data itself - or not, depending on mutability and size).
Regardless, none of that changes the fact that it's a bad design decision, because people regularly get it wrong.
As a non-Java developer I have to ask why you think this is a reasonable design choice. It's a massive footgun for the a casual polyglot Java developer. I can imagine just forgetting I can't do == if I've just switched from being immersed in another language.
Which is surely the important bit. I can just use "==" in C#.
I'm parsing your comment to mean, "it's easier to write correct code in compiled languages", but this is not obvious to me, or anyone who's, for instance, written any C at all.
* Ahead of time compiled languages reduce flexibility but allow the compiler to do more reasoning about the system.
* Statically typed languages allow the compiler to reason about the assignments you make and methods you call. If the language is also AOT compiled, you avoid crashing at runtime.
* Dynamically typed languages significantly reduce boilerplate and avoid the mental effort of expressing an idea in a rigorously typed manner. It's the "hold my beer" approach.
There's a usual confusion of banking in this case, but if we're talking about popular dynamic/memory-managed languages, you can put an equivalent of:
try:
actual code
except:
log.exception
At a high level and be fairly sure things will be ok in a long run even with failures along the way.On the other hand, most static-typed systems handle failures explicitly.
I think C is a bad example here. We've come a long way since C. We know how to do better.
Unless you want to include the ecosystem as well. C + valgrind + PVS + clang-analyser do make it easier to create correct code.
Point is, code culture matters much more than language.
The advantage of this is that if you end up not wasting too much time "building the wrong thing". Let's say that you took one form of I/O and built massively strict validation in and then realized later that you should have taken an entirely different form of I/O for your subsystem. All that time building in validation on that useless part of code was a pointless waste.
I don't have any stats, but my gut feel is that on average 40% of code can end up being tossed in this way (in some projects it's 100% =).
Prototyping speed is, additionally, not just useful in reducing the cost of building the right kind of code, it's useful in reducing the cost of building the right kind of test (a really underappreciated facet of building mission critical systems).
In my younger years I used to believe that for mission critical systems "building the wrong thing" was somehow less of a problem in code because you could fix requirements and do architecture upfront with some sort of genius architect. Turns out this was wrong.
I'm curious as to where you have observed this, because my experience has been exactly the opposite: even in circumstances where "dynamic" might lead to more readable code, C# developers are loathe to use it, to the point where it's very hard to find it in idiomatic C# code.
(1) Due to ecosystem richness, the need to collaborate cross-functionally, and other reasons.
EDIT: grammar gremlins.
If you don't believe the results of the study, please provide the counter evidence.
Besides, all I wanted to do was to point out that you shouldn’t derive any conclusions from that study. I don’t need any counter evidence for that (apart from the evidence that the study is flawed of course).
Instead, common sense is actually the best substitute we have.
Alternatively, usage of Rust may not be correlated with fewer bugs overall but I'd expect there to be fewer of the use-after-free variety.
I actually heard Aribus software is written in C++. If true it should scare the shit out of anyone.
In practice I find that static typing as done by, say, Java and C#, excels at uncovering very shallow bugs. Its benefit is the most obvious when you do very little automated testing. The tech industry's dirty little secret is that automated testing is done infrequently and badly, so it looks pretty effective.
I find that if you try to compare the cost of building sophisticated tests in a high productivity language (like python) with the cost of writing code in a low productivity/high type safety language like haskell, there's a pretty blurry cost/benefit trade off. If you take out productivity as a factor, haskell looks much better.
Nonetheless, the idea that static types (and, to a lesser extent, formal proofs) are some sort of silver bullet persists. There is no silver bullet.
On that note, why would you consider Haskell a low productivity language? People who are well versed in it seem to find it an exceptionally productive language to work with.
Types also help document code. Just like tests.
The point is that both are investments and both have different payoff matrices. Sophisticated are often better at preventing obscure logical bugs. They're also good at uncovering obscure not-bugs and preventing code from getting out until the compiler is satisfied. Bad unit tests also do this.
Haskell simply takes longer to write than other languages. Given two developers of equal skill and experience, anyway. I partly attribute the relative paucity of haskell software out there to this.
Further, while I agree with your “different matrixes of payout”, I think static types are a very low investment and they have a respectable payout in terms of preventing bugs but also in terms of documenting code and facilitating tooling (such as autocomplete or documentation generation) and they also permit easier code changes than comprehensive unit testing (even “good” unit testing). Of course, I’m not advocating that static types completely obviate tests; rather that they obviate some of the tests; however, there is no clear answer as to how many or which tests are obviated; it’s very circumstantial.
> Skepticism is not pessimism, however. Although we see no startling breakthroughs, and indeed, believe such to be inconsistent with the nature of software, many encouraging innovations are under way. A disciplined, consistent effort to develop, propagate, and exploit them should indeed yield an order-of-magnitude improvement. There is no royal road, but there is a road.
A good counterpoint is Go. Despite being far more familiar with Python (it’s my professional/work language), it’s static types actually help me write correct code more quickly than I can with Python. Your mileage may vary, but I think anyone who has given Go an honest shake will find that it’s at least in the same productivity ballpark as Python. Note that this isn’t considering tooling or deployment, or performance where Go specifically excels over Python (and also often due to static types).
Of course, with sufficient investment in testing, you could get the same confidence with Python that you get in static languages, but writing tests is a productivity cost as much as boilerplate or a pedantic compiler. And with static languages, I often find that I can write fewer tests for the same confidence (in fact, I often prototype in Go and backport to Python).
Before test suites, I would have agreed with you.
The biggest value to me is that you can easily do huge refactorings with a lot of confidence. That in turn means you can keep redesigning your code and frameworks long after they would ossify into legacy code no one dares change in a normal project.
That's not to say I disagree that much with your point.
Or rather, how would you define mission critical?
You realize that many mission-critical workloads run on Python with great success, right?
https://stackoverflow.com/questions/3270680/how-does-python-...
"CPython implementation detail: Objects of different types except numbers are ordered by their type names; objects of the same types that don’t support proper comparison are ordered by their address."
I wouldn't say "100" < "2" really counts, since, alphabetically it makes sense (an equivalent would be "baa" < "c").
Seems like a completely reasonable thing to do?
????????
None (as far as I remember) of the common "wat" examples about JS are interpreter implementation details, but specified behavior.
I'm mildly curious why CPython chooses to implement it this way, but if I had to guess: I'm assuming it is to provide a stable order between objects of different classes with the same hash() value. Hash-value being what I suspect is what the quoted answer misrepresents as "their address" (the address being in CPython the default fallback for objects that do not implement a hash() function), and being a good candidate to establish an arbitrary order.
While static typing does catch code errors, in my experience they are mostly errors that would be caught somewhere else in the testing process. Memory safety issues are rarer, but also have a penchant for showing up unexpectedly and for crashing the entire system in unrecoverable ways.
Obviously, the best would be to have both kinds of safety, but people program safety critical systems in C++ all the time, and deal with the consequences. At least you can catch a TypeError in Python and handle it.
Catching TypeError in python and handling it is not a substitute for static typing.
Realising that type checking is a good thing, then bolting it on, I feel is a crappy solution, when you should just use something designed better from the start.
I completely disagree that there is a "productivity toll". If you dont know how to speak a language then obviously it seems difficult. I think you save time in the long run not dealing with all the type errors.
Have you written and deployed production code in the languages you mention above? Did you encounter a substantial number of type errors?
Is a '.close()' method being called twice a type error? It is in a language that can express states in the type system.
Is SQL injection a type error? It is when you use refined types in your sql library interface.
The vast, vast majority of errors I run into are errors that, with effort and the right type system, I could turn into type errors.
The point being that asking "would those be type errors?" is a really big question that is, when answered simply, "yes".
Pyflakes and unit tests will detect the great majority of potential errors. Mypy gets you closer to zero.
No, it's a pragmatic argument. There's no circularity it.
If they're not mainstream, then they're neither "very usable" or "well-supported" to any extend that a mainstream language would be.
- There are no mainstream languages with powerful and convenient type systems
- But there are less mainstream ones
- But I won't use those
- There are no mainstream languages with powerful and convenient type systems
That's however a totally legitimate engineering choice.
Engineers picking on a language shouldn't bet on less mainstream incomplete environments with less tooling and libs and options and devs, just so that they can raise them into the mainstream.
That might be a "tragedy of the commons" thing, but it's not an engineering obligation to be an early adopter.
> There aren't any mainstream languages that have the ability to detect the ultra high-level errors you describe.
And it's like, well there are great languages that can do that, but if you restrict yourself to that tiny 'mainstream' subset, you'll never know it.
And moreover I think engineers (or rather, companies) are way too conservative about this stuff–picking 'mainstream' tech can be a touch-and-go proposition at any time. Just because something is mainstream, doesn't mean it's the right choice for your project.
Injections are handled by any language with taint, although that became less popular. Perl had it. Anything with typed orm also has something similar (for example Esqueleto in Haskell)
As for closing resources, you can use a pattern such as using(open("file.txt"), (f) => { ... }) if your language doesn't already support such a construct.
Sure, no one should need to do that anymore, it's just an example. There are many other cases that are similar, however.
`using` is a language construct, not something that is part of an interface. It also does not prevent close from being called twice.
It's when you want to add features to a codebase you don't understand, when these features shine.
TypeScript's type system is far more powerful than Java's but I'm not sure that's a good thing. A lot of the concepts in TypeScript are necessary to paper over gratuitously dynamic APIs in JavaScript.
While Java certainly has warts in its type system (no generics over primitives, arrays), Java's type system straight-forward compared to TypeScript. TypeScript stuffs a staggering amount of functionality into the type system. To name a few of the more esoteric features: literal types, const assertions, asserts modifier for type predicates.
Regarding your criticism that `result.ok()` returns `None` if the value is not an `Ok` type: That's what `result.unwrap()` and `result.expect(msg)` are for :) (There's no unwrap_err and expect_err so far, but PRs are welcome.) The lib is strongly inspired by Rust (see https://doc.rust-lang.org/std/result/enum.Result.html), that's why `result.ok()` returns something option-ish.
I like your `isinstance` approach btw! Will have to think about it a bit more. Reminds me a bit of Typescript as well. Does mypy have type guards (https://www.typescriptlang.org/docs/handbook/advanced-types....)? Maybe that could be used to create helper functions that check for success/failure and which help mypy to derive the correct type.
not yet, maybe in the future. might be possible to emulate some use cases with Literal Types. tracking issue:
Now to sneak it in at work...
Also doing lots of messing with data for personal stuff (e.g. https://beepb00p.xyz/mypkg.html#examples). But yep, at least for me, the reality is if I need some quick yet useful code, Python happens to get me there very fast, just because of the sheer amount of code people already have written in Python.
Feel free to give it a shot! https://pypi.org/project/safetywrap/
I think naming the latter "Encode Typesystem" is misleading (and your question is an indication that I am right).
its a very common name.
T1 = TypeVar("T1")
T2 = TypeVar("T2")
GenericUnion = Union[T1, T2] def foo() -> OneOf[Bar, Baz]:
if blah():
return Bar()
return Baz()
def bar_to_blah(bar:Bar) -> Blah :
...
def baz_to_blah(baz:Baz) -> Blah :
...
one_of_bar_or_baz = foo()
blah = one_of_bar_or_baz.match(bar_to_blah, baz_to_blah)
but the critical thing is to get a type error if you add another case to foo's return type, or change the generics to different types.