Python types have an expectations problem
medium.com
medium.com
In Python, yeah, they're called Type _Hints_ for a reason. Don't count on them at runtime here either.
Both are dynamic languages, it's hard doing anything meta/schema driven with rigid types. If you really want a hard type system, just move to GO or Rust or C or something with a real type system enforced.
#include <stdio.h>
int main(void) {
unsigned int positive_number = -1;
printf("%d", positive_number);
return 0;
}
Prints -1.And printing it as %d is technically a misuse of printf, since %d means print as a signed integer. If you did %u (print as unsigned integer) instead you'd see the value is really 4,294,967,295 (again, platform dependant)
So in C if I assign a char literal to an int, I get an int. If I do the same in Python, I get a string.
In a strongly typed language, which I'm defining as "needs the left and right hand side of an assignment expression to match somehow", I'd get a type error.
I'm not sure I understand - in what situation would you need runtime type checking? I always assumed the situation was similar to Haskell where the types are erased at runtime but it's still impossible (absent opting in to unsafe stuff) to have type errors at runtime.
You NEED to verify data from outside your app anyways but if you have runtime checking, trying to cheat and skip that step is harder.
There's also a myriad of ways to alter data running in a browser. From plugins to just fooling around with the console in DevTools.
On a large enough team, you might not be able to trust the code calling your code. It's easy to slip 'as whatever' in there. I do it when using vue's reactive objects instead of fighting with typescript.
> Don't count on them at runtime here either
If you're adventurous enough, you can reflect on the type hints and check things yourself at runtime, but you have to understand that the type hints aren't meant for this and they could well blow up in your face. Still, I've had some success constructing dataclass instances from JSON objects based on what fields() tells me about the attribute types. Whether you want to do this yourself in production depends on your tolerance for hilarious edge cases and interpreters that get to do things differently because nobody promised you anything.
So is it possible for me to enforce type hints at runtime, so my program crashes if a type is ever wrong? How do I turn this checking on?
That said, using these sorts of tools (or trying to enable runtime type checks in general) is usually a misunderstanding of how the type checking works. Think of type annotations less as a way to enforce types, and more as a way to explicitly declare intention. If you annotate a function as taking a string and returning an int, then you're saying that, if the caller obeys their side of the deal and passes you a string, then you'll obey your side of the deal and return an integer. Moreover, a type checker can demonstrate that this is the case: that for every path through your function, assuming the input is always a string, then you will return an integer. This frees you from runtime checks entirely, because the checks can happen entirely at compile time: assuming the input types are correct, your code will return values of the correct type in return. If your whole program is type checked, then we can be confident that the whole program behaves correctly in terms of types.
There are two big exceptions to this. Firstly, runtime data cannot be checked in this way, so needs to be checked at runtime. Usually this happens at very explicit boundaries: you load a TOML file, and use Pydantic to validate the file and parse it into a Python structure. If the Pydantic check succeeded, then you know the data had the correct types, and therefore that your function will behave correctly as outlined above. If it failed, you get a runtime error at the exact point that the invalid data enters your system. The danger is doing something like json.loads and then not validating the result in any way.
The second exception is when you manually choose to ignore the type system and tell it what's going on. This includes type assertions where you override Mypy's expectations for what a type should be, or the use of the `Any` type. In these situations, you deliberately can the compact you've made with the type checker, but you also get a chance to handle cases of extreme dynamism in Python that the type checker can't deal with. These sections of your code will naturally be unsafe, but are usually more explicit and so can be checked more carefully in code review.
The result is that, with a type system that handles this stuff well, you do not need to enforce type hints at runtime (and trying to enforce type hints at runtime is usually a sign of not leaning sufficiently into the type system in the first place). That said, "a type system that handles this stuff well" is the key phrase here - Typescript is a good example of a tool that works well as a static-only type checker, in that it gives you a lot of power to define types that match the underlying runtime. In my experience, there isn't yet a tool powerful enough to handle that in Python, which can make developing with type hints quite frustrating at times. But I don't think runtime type checking would change that at all.
See for example https://pypi.org/project/py-strict-typing/
Add static type checking of your code through mypy or pyright, and now you can reasonably guarantee that your code is type safe, and stuff interacting with your code is type safe.
Not perfect, but it's a reasonable approach for important libraries!
> If you really want a hard type system, just move to GO or Rust or C or something with a real type system enforced.
For what it's worth, Rust also has no types at runtime, which is why Rust doesn't have reflection and relies heavily on compile-time magic with macros. The difference is that there's no way to get a binary to run if the code doesn't typecheck. The most underrated feature of a compiler is that it can say "no"; in the case of Rust, the killer feature of the language isn't what you're allowed to do, but what you're _not_ allowed to do, even accidentally.
And this isn't just a side effect of the build tools, but a genuinely useful feature. While I'm working on a change, I can run half-finished code and get some feedback before fixing the rest. This is particularly useful for tests - I can test a module, even if I've not yet updated all of that module's usages, to sanity check whether what I'm doing makes sense.
There can definitely be downsides to this separation of type system and runtime behaviour, but it's also very useful, and it works the same in Typescript and Python.
I think the biggest issue that Python's type checking has in comparison to Typescript is just sheer power: Typescript can very explicitly and correctly type real-world Javascript code, whereas typed Python, in my experience, ends up feeling a lot more like old-school Java than idiomatic Python. And it's difficult to sell old-school Java to Python developers.
I think the difference is really one of UX (or DX I guess) and availability: Yes, if you already work with a toolchain, there won't be a difference. However, lots of people don't. If you just use the language runtime on its own and want to get your program running with the minimum possible effort, then for typescript, you'll run the transpiler and feed the resulting JS to a browser or nodejs vm; for python, you'll just run your program with the python interpreter directly.
The thing is that for the "minimum" typescript workflow, type checking is per default performed, while for the minimum python workflow it isn't. That you could also disable type checking for typescript or use a linter to get it done in python is besides the point - those are additional options and tools that you have to spend additional effort to activate - and you have to know about them in the first place. Someone who just has some basic skills in the language is unlikely to do so, so effectively, type checking for them is performed in typescript but not in python.
So there isn't really a "minimum" Typescript workflow because by using Typescript, you're already choosing to install additional tools and set up a more complicated workflow (compile then run, as opposed to just running Javascript). And if you'd do that with Javascript, you'd do that with Typescript as well.
(I think it's no coincidence that the long-term plan from the Typescript team is to have type annotations become part of regular Javascript syntax, in the same way that Python includes type annotations out of the box. In both cases, they'd be ignored by the parser and only used as metadata for the purposes of a separate type checking tool. The Typescript team do not see themselves as developing a separate language to Javascript, but rather just a dialect that includes types.)
Like:
somevar: dict[int, int] = {0: 0, 1: 1}
Produces "TypeError: 'type' object is not subscriptable"Hard-to-grok errors tend to appear for every language and every backwards-compatible change unless the language itself specifies the target version in the source code file itself. The only language I know is doing this is Solidity.
I wanted to check if the use case can be generalized for all situations. For example, the code below will throw a runtime error
* Input should be a valid integer, unable to parse string as an integer [type=int_parsing, input_value='A', input_type=str] *
```
from pydantic import BaseModel
class User(BaseModel):
id: int
name: str
def print_user(user: User):
print(user.name, user.id)
user_1 = User(id="A", name="John")
print_user(user_1)
```I have also been using typescript with React Application and really find it better compared to python due to type checking. But I can still imagine a future when type-checking gets incorporated in native python, unlike javascript/typescript. The groundwork has been laid already with type hints.
For example:
def fooify(x):
if isinstance(x, list):
map(fooify, x)
else:
x * 2
I guess this is to make scripting more forgiving.In typical typed languages, you would have two functions instead:
fooify : number -> number
fooifyMany : [number] -> [number]
But in the Python community, it's common to have a big function with many behaviors. However, then the type annotations cannot be so precise: fooify : any -> any
Not an experienced Python dev, so curious how this works in practice.def fooify(x: int | list[int])
If x is a list, the result is a list
I don’t think that the type system can describe this.
In fact, I think TypeScript will, for your given example, with the 'else' clause accurately identify the type of x to be a float if it's a union like the one I wrote down above.
In fact, in a lot of languages, you can't tell this is a float, statically. It would work with whatever type was passed in that supports multiplication. And return the appropriate type. Right?
Using you example:
from typing import overload
@overload
def fooify(x: int) -> int:
...
@overload
def fooify(x: list[int]) -> list[int]:
...
def fooify(x: list[int] | int) -> list[int] | int:
if isinstance(x, list):
return [fooify(_x) for _x in x]
return x * 2
[0] https://docs.python.org/3/library/typing.html#overload T = TypeVar("T", bound=int | list[int])
def(x: T) -> T:Meanwhile it's not even possible to express such things in other static type systems. So I'm not exactly an unhappy customer, but it does put certain things tantalizingly close, but still out of reach without a ton of clunky boilerplate and LoC explosion.
what do you mean? It seems relatively straightforward.
type CompoundInt = int | list[CompoundInt]
def fooify(x: CompoundInt) -> CompoundInt:
if isinstance(x, int):
return x * 2
else:
return list(map(fooify, x))
print(fooify([1, [3, 4, 5], [6, 7, 8], [[[4, 4, 4]]]]))
This uses the `type` keyword introduced in 3.12. Unfortunately Mypy doesn't support it yet :( so this workaround can be used instead: from __future__ import annotations
from typing import TYPE_CHECKING
if TYPE_CHECKING:
CompoundInt = int | list[CompoundInt]Instead of this:
def check_permission(user: User, perm: str, obj: BaseModel | None) -> bool:
I find it much nicer to read something this: """
Check if user is allowed to drive a car
:user: User # The application's User model
:perm: str # Our magic permission string. Tab seperated.
:obj : Basemodel | None # Base model of all our objects since 2017
:returns bool # True if the user is allowed to drive
"""
def check_permission(user, perm, obj):
This way I can grok the code much faster and only have to look into the type declaration when I want to. Due to syntax highlighting, it will look really nice. Because the whole type part can be styled in a dimmer color which puts it into the background. And I can define a key combo to show/hide the whole type part.def extract_coordiates( unreseted_idx_from_df, json_data: dict, notna_idxs: list, na_idxs: list ) -> list:
"""
Extracting coordinates from the json_data['data'] and returning the 4
positions of those as an array
Parameters
----------
json_data: json
Data extracted from tabula.read_pdf and using output_format='json'
notna_idxs: list
List of id's from the dataframe that ha no na values.
na_idxs: list
List of id's from the dataframe with na
Returns
-------
List of tuple coordinates:
[(1, 2), (2, 2), (4, 3), (1, 4)]
"""
Thing is, that you may want to get fast information in the linter with this type hinting. If you need to read the documentation on big functions over and over, that means that's not clear at all, and having basic type hinting while typing the atributes is going to be clearer.You are right, that some docstrings are necesary. But that does not mean that is the best practice. The best practice is use both, type hinting and docstrings.
Yes. That's how people successfully use type checking in Python.
Trivial to set up but ends up being a huge time saver, especially when reviewing junior colleagues code - don't even ping me to review your code until you've managed to convince the static analyzer it will work!
But after that they learn to write code which makes both the static analyzer and the reviewer happy, and such code tends to be much more maintainable down the line.
When you don't have strict typing discipline enforced by the CI, you will likely have to enforce it manually anyway because projects in a gradually typed language without simple types enforced tend to become a complete mess IMO.
It installs shims around all your methods which type check arguments on the way in, and results on the way out. Of course this comes with a performance penalty, so it’s often enabled for development and testing, but disabled in production.
This issue compounds in a painful way. Because 99% of your codebase is starting out untyped, you have a couple of options, neither of which I have found to be practically very useful.
For the first option, you can run a blanket `mypy` invocation on your entire codebase and have a massive blast of errors you ignore for some time. Because it's necessarily going to error in the beginning, you can't really fail your CI pipeline as a result of this yet. If you can convince your team to gradually improve typing or set a deadline for eventual CI failure based on types, you might be able to move the needle and eventually get your codebase typed.
From my experience though, this basically just became a CI step everyone ignored, and for people on the team not passionate about typing, they never worried about it.
The other option, which is far more annoying in practice, is to pick a few "seed" files in your codebase that you can add typing information to quickly. Then, supplement your `mypy` invocation with a list of these seed files so it becomes something like `mypy file1.py file2.py ...`
As you continue to improve typing "at the edges" of your codebase, you gradually add more and more files to the `mypy` invocation until you're eventually (hopefully) adding entire subfolders, and then maybe eventually the entire codebase. Starting at the edges means you can enforce the CI check from the beginning and get value quickly.
The issue here is mostly remembering to continue to add files to the `mypy` invocation, which means you're constantly altering your CI pipeline. You know how when you alter a CI command it breaks sometimes because you got the encantation slightly wrong? Multiply this effect across basically every member of your team 1x a week because most people probably haven't edited your CI pipeline before. With even a small team (~7-10) making constant changes to a codebase, this quickly becomes extremely painful, and pipeline failures start eating a significant chunk of time just trying to debug if the encantation is wrong or if the types are actually broken.
We mitigated this by having only one dev add new files to the `mypy` invocation, which worked well for the CI side of the story.
The local side of the story is what ultimately led to enough fatigue to give up. It was hard to get in the rhythm of using local `mypy ...` invocations to check your types as you made changes, and so the experience for most of our team was to push changes, and then the types would break in CI, which was frustrating. They'd go in and try to fix it, and sometimes Python typing gets weird, and a fix wasn't immediately obvious. Eventually you get to `#type: ignore` or `Any`s being thrown around to sidestep the CI pipeline, and your typing story has collapsed again. The real kicker for us was the painful juxtaposition between `mypy` and `import`s. Is the giant swath of errors I'm seeing from this file or from a file I imported? Asking the entire team to become Python typing gurus to sort out these issues was a non-starter.
Does anyone have experience gradually adopting Python typing in a large, older codebase successfully? If so, would you mind sharing the methodology you found success with?
And yea, from an IDE perspective, it seems like maybe a sensible default would be to run `mypy ${CURRENT_FILE}` on save or something -- I've tried this manually myself, and it's decent, but it's not good enough to use constantly. I don't remember specific issues since I haven't done this in a long time.
Pre-commits are painful (on purpose!), but preventing people from shipping seemed like the wrong call. So we ended up dropping them.
Many people, including myself, also like making atomic commits that may not be fully typed yet. Conditioning people to skip the hook(s), doubly so during a gradual typing rollout, just conditions them to skip them always. Leave it in CI where it's checked at a point in time where it actually matters.
Rather than doing this, which does indeed seem like a headache, it may make more sense to skip import following at the very beginning until your core is typed so you can still enforce typing on the leaf nodes moving forward.
> Eventually you get to `#type: ignore` or `Any`s being thrown around to sidestep the CI pipeline, and your typing story has collapsed again
While there are some cases where this is truly the best option, ultimately you get to the point where you just don't allow this, otherwise what's the point of all the effort?
> and for people on the team not passionate about typing, they never worried about it <...> Asking the entire team to become Python typing gurus to sort out these issues was a non-starter.
The faster the core can be typed (and typed correctly), the easier it becomes for those who are less passionate. Presumably someone has done the calculus to determine that this effort is worthwhile, so while the team doesn't necessarily all have to reach guru level, they need to be convinced to continue the work. Removing barriers is huge for this, since as you've noticed once it starts being easy to ignore it's really challenging to stop ignoring.
Yea, this is solid advice and something we did at one point. It's essentially strictly necessary for an older codebase.
> While there are some cases where this is truly the best option, ultimately you get to the point where you just don't allow this, otherwise what's the point of all the effort?
Again, true! We reached fatigue and gave up long before we hit the point where this would have mattered.
> The faster the core can be typed (and typed correctly), the easier it becomes for those who are less passionate. Presumably someone has done the calculus to determine that this effort is worthwhile, so while the team doesn't necessarily all have to reach guru level, they need to be convinced to continue the work. Removing barriers is huge for this, since as you've noticed once it starts being easy to ignore it's really challenging to stop ignoring.
I think this was probably my biggest failure in terms of the success of adding typing. I severely underestimated how long it would take to add typing info to some of the really old pieces of code, and it wasn't reasonable to expect to be able just to sit down and add types without delivering business value for an extended period of time.
My inexperience with `mypy` and the general typing ecosystem in python contributed significantly to the team ultimately reaching fatigue and deciding to give up on it (mostly).
See https://mypy.readthedocs.io/en/stable/existing_code.html for some more advice.
I was pretty green to the Python typing ecosystem when I started implementing it in our large pre-existing codebase, and I did not lean heavily enough on a `mypy` configuration that was module-specific.
It would have saved me a significant headache, and in hindsight, this seems like the main viable option for typing an old codebase effectively. This gets extremely gross when you have 300 or 400+ submodules across your codebase, but start small and work your way from the outside in if you want the best chance of success.
Are you converting bespoke dicts to sensible NamedTuples/dataclasses? Or are you purely adding type hints?
It's nice to be able to hit `.` in your editor and have the options for whatever object you're staring at pop into a list you can pick from. Similarly, for a `TypedDict`, you can hit `["` and just have all the possible key options autocomplete.
The bespoke dicts mostly come from early-days SQL queries akin to
SELECT * FROM single_table
I wrote a small tool in Rust (yes, I was looking for an excuse to use Rust at work) first to create a giant `TypedDict` file that basically typed all the table rows in our primary database and then a small parser that reads our codebase for these trivial select * from single_table queries and adds `TypedDict` hinting for autocompletion purposes going forward.So a line like:
some_result = curs.fetchone()
becomes: some_result: SingleTable = curs.fetchone()
Then when you type: some_result["
You get a lovely list of auto-completed possible keys instead of having to go hunting through the table to remember what the column was called.The other reasons we wanted to implement typing were all the main reasons you'd want a typed codebase in the first place. Shift a lot of mistakes to build-time instead of deployed-in-production time.
Static typing significantly increases code length compared to duck typing, significantly increases development times and significantly increases bug count per delivered software feature.
It additionally gives developers the false impression that you can write large projects in Python without splitting the codebase into small micro-services. Hint: You can't, it is always a complete disaster as Python's support for static typing isn't good enough for large projects.
Going to need a source for this.
You can easily find information on duck typed programs being significantly shorter and that statically typed and ducked typed programs have the same bug count per line. You can also look at the language defective tables showing ducked type Clojure as having the least bugs in commerical software and statically typed C++ having the most bugs in commerical software.
Google is your friend here.
"Number of bugs will drastically increase" is simply wrong.
Type checking prevents bugs before they ship.
Type checking mostly shows that code can be compiled successfully. The thing it was designed to do.
i've read some of these studies, they don't prove much. it's typically languages with inexpressive type systems like java.
>Type checking mostly shows that code can be compiled successfully. The thing it was designed to do.
what languages are you thinking of when you say things like this? these debates are useless because one person is thinking of the type system of language X and the other person is thinking of language Y. a lot of heat but little light.
I don't see how e.g. [1] or [2] are "falling flat on their faces".
[1] https://fsharpforfunandprofit.com/posts/designing-with-types...
[2] https://lexi-lambda.github.io/blog/2019/11/05/parse-don-t-va...
or even something as basic as non-nullable types, enforced at compile time. a clear, obvious win.
The duck typed code is also significantly shorter, less lines of code per software feature means less bugs per released software feature.
It's very difficult to compete against just having less code. You only needed to write 1000 lines of code instead of 3000 lines of code? Then you just eliminated 66% of the bugs.
Static typing on the other hand catches less than 1% of bugs.
So you can go for the duck typing approach and remove 66% of the bugs in your code or the static typing approach and remove 1% of the bugs in your code.
Which approach do you think works better?
If you're not going to bring bold numbers to support bold claims, why should any of us bother listening?
I mean you could always try to find some sources to try and prove me wrong if you want?
"I mean you could always try to find some sources to try and prove me wrong if you want? "
That's not really how it works though. If you make a claim, either you back-it up or it's irrelevant.
Case and point: I believe that a flying tea pot created the universe. By your logic I am right and it should be up to you to prove me wrong.
So to conclude this argument, if you are correct, you should prove it as the person who responded to you asked you to do.
Failing that,your comment goes back to being nothing more than your opinion instead of a fact.
If s/he thinks that what I'm saying is wrong so strongly then they can go back that up with something. They will obviously fail terribly because I'm correct but that is their choice.
I use python add my day job, and I hate writing and dealing with typing. I do admit it makes it easier to reason about other people's code though, especially when faced with libraries that don't have any hints.
Docstrings are usually a lot better than type hints for that purpose.
I could write "Humans have actually always had blue skin and hair, but due to a cosmic ray warping our ocular nervous response systems 1000 years ago, we now see humans the way we see them today". And I'd damn well hope someone would ask me to back this absurdity up with evidence.
... and thus: "lalalala I can't hear youuuuuuuuu over the sound of being right!" doesn't make for a good look, is basically what I'm telling you, and asking you to please either cite some sources, or stop doubling down on having to be right because you said so. This is such a ridiculous exchange I'd swear I was having it as an elementary school teacher.
The error won't exactly point me to the source of bug and I was getting vague message that some data was missing .
And there were many instances of same throughout my project experience.
With typescript you will know instantly. The javascript being weakly typed compared to python necessitates use of typescript.
I can say from that experience that pretty much everything you are claiming is wrong. Developers were enthusiastic about adopting type annotations. We found it made code easier to understand, gave better support from IDEs, made refactors easier, and caught bugs earlier.
Type-checking that codebase with Mypy was a technical challenge, but the Mypy team did a tremendous amount of work scale Mypy through caching and other optimizations. It was still slow, taking minutes at times, but way faster than the complete test suite.
Now, I will be the first to tell you not to write a 2Mloc server in Python, but if you happen to have created one, type annotations are huge boon to making it work.
I'm telling you the ducked typed micro-service based approach is the way to go.
But you sort of know that "not to write a 2Mloc server in Python", don't you? If we encourage developers to use static typing they will be writing 2Mloc monolithic servers. That's just the type of code that static typing encourages.
Encouraging static typing in Python is putting them down the path of disaster.
"If we encourage developers to use static typing they will be writing 2Mloc monolithic servers."
I don't even know where to begin with this wild assertion. Ill-advised as it was, Dropbox wrote its two million line server long before typed Python even existed. I wasn't around when they started cracking it up (see blog post below), but I strongly suspect the type annotations in the codebase helped rather than hindered that effort.
https://dropbox.tech/infrastructure/atlas--our-journey-from-...
Sure, if one doesn't care about their code quality, after all everyone loves to write unit tests for every possible use case.
Static type checking has a measured bug catch rate of <1%. So the quality of any code where the programmer depends on static typing to verify correctness rather than unit tests is very very low.
Looking forward to a set of paper from renowned researchers with SIGPLAN/ACM, IEEE proven credentials in language research.
The relevant term is Type Checker. Static Type Checker: Something that checks types statically (i.e. by inspecting the source code without running it). Doesn't necessarily say anything about compilation, execution strategy, or indeed if the user program will ever be run at all.
The majority of bugs in code are about the behavior of the code rather than the typing. You require unit tests or some other more advanced technique.
“Python type hints are a core part of the language, they even have standard library modules (typing), and yet they don’t do anything when used in that language without some external tooling.
That, to me, is a bit of an expectations mismatch.”
I.e. they’re complaining that a type checker like mypy isn’t run by default.
Let's think of this in a deploy environment like kubernetes -- if you don't have such a check, then you could deploy code that fails _at runtime_, causing an outage, because someone made a mistake with types.
If you fail _at startup_, then the deploy will never go healthy and will fail, leaving the old pods still running, causing no interruption in service.
And yes you _should_ have this check in CI, but there's no reason not to have defense in depth.
Also, the comparison to typescript is not quite fair because you can't run Typescript in your browser; you have to set up a proper build system. You can do the same thing for Python and make typechecking (and linting and code formatting) part of your build and you get the exact same benefits.