Structural pattern matching in Python 3.10
benhoyt.com
benhoyt.com
>A related feature to final classes would be Scala-style sealed classes, where a class is allowed to be inherited only by classes defined in the same module. Sealed classes seem most useful in combination with pattern matching, so it does not seem to justify the complexity in our case. This could be revisited in the future.
Has this been picked up? Replacing long if-elif blocks is only half the power of pattern matching. The other is to have exhaustiveness-checking by the static type checker.
I have pieces of code that would really benefit from this, and I was looking forward to migrating it to the new pattern matching feature. But I can't find any word on whether this is being revisited. That's disappointing.
class Final(type):
def __new__(cls, name, bases, dct):
if any(isinstance(base, Final) for base in bases):
raise TypeError("No!")
return type.__new__(cls, name, bases, dct)
class NoSubclasses(metaclass=Final):
pass
class WillThrow(NoSubclasses):
pass from enum import Enum
from typing import NoReturn
class Stuff(Enum):
Cat = 1
Dog = 2
Bear = 3
def exhaustiveness_check(value: NoReturn) -> NoReturn:
raise AssertionError()
def do_thing(stuff: Stuff) -> None:
if stuff is Stuff.Cat:
print("meow")
elif stuff is Stuff.Dog:
print("woof")
else:
exhaustiveness_check(stuff)
And if you add a check for Bear it will validate it!https://mail.python.org/archives/list/typing-sig@python.org/...
I have a fork of the "adt" library that supports sealed classes. For static type checking, I'm using a fork of typpete, a SMT based type checker.
It's not as user friendly as mypy or pyright, but in the longer term can be a winner because it supports user defined type constraints (aka refinement types).
I remember that in the early 00's people called it "executable pseudocode". I bet nobody would use now that expression for modern Python with decorators, comprehensions, walruses and the like.
Python doesn't need more features but better implementations, IMO.
Edit: Just remembered the Zen of Python, I don't think modern Python is following it.
...
Simple is better than complex.
...
Readability counts.
...
There should be one-- and preferably only one --obvious way to do it.
Having one way of doing something and sticking with it even when better ways emerge is nonsensical imo.
f-strings are more readable (for me) and concise than string formatting or .format()
I do agree on the implementation front though, many of the other implementations are still on 2.7 or 3.6 at most, and don't have mainstream appeal. Made more with interop in mind (like ironpython with .net and jython with the jvm)
Not for beginners though. One of Pythons main strengths was that it looked simple. Not anymore.
Point is Python used to be a complete yet simple language with few constructs. It more or less followed the path of least surprise and people revered it for that. For several years, it feels like PEPs just add more and more clutter to that foundation.
It sort of looks like an arms race and I just don't get why or what's going on. Certainly, Python doesn't feel the same anymore.
elif cmd.op == 'TUPLE':
arity, = cmd.args
tuple_args = []
for _ in range(arity):
tuple_args.append(self.data_stack.pop())
result = '{' + ', '.join(map(str, tuple_args)) + '}'
self.data_stack.append(result)
Is following use of comprehension any better? I don't think so: elif cmd.op == 'TUPLE':
arity, = cmd.args
tuple_args = [str(arg) for arg in reversed(self.data_stack[-arity:])]
self.data_stack = self.data_stack[:-arity]
result = '{' + ', '.join(map(str, tuple_args)) + '}'
self.data_stack.append(result)
Any other suggestions for improvement?There's usually a threshold where I go "nope" and unroll it into a good old loop.
But when you're in their sweet spot, they definitely improve readability.
elif cmd.op == 'TUPLE':
arity, = cmd.args
stack = self.data_stack
text = ', '.join(str(stack.pop())
for _ in range(arity))
stack.append(f'{{{text}}}')Buy I personally like it when I read other people's code and see something which I don't understand, which I've never seen, and then it turns out to be a feature which actually makes coding easier and more efficient. Then I think about all the places where I could have used that feature in my old code.
Code examples in docs and tutorials are still easy, libs api still clean, code base still readable.
Optional things stayed indeed, optional.
My grip is actually the contrary, many python code bases are stuck in the past. But I'm ok with it, people are not all pro devs, they need a productive language subset.
And the new syntax is providing some fantastic new tools, such as pydantic/fastapi/typer (if you haven't try it, it's really paradigm shift, and I say that as a django fan). f-string are awesome. dataclasses are nothing but convenient and readable. Advanced unpacking as well. I seldom use the walrus, almost nobody does, but when one does, it's useful.
> decorators, comprehensions,
Those existed in python 2, 10 years ago. It's hardly new python. And web frameworks would be terrible without those.
From the outside looking in the Python community seems aggressively anti reworking their code to adapt to new language features. The 2 to 3 debacle comes to mind when compared to other languages, like JS, which made a much bigger change and everyone jumped in it and rewrote their code to adapt, or Ruby which went through some backwards incompatible changes at the same time as Python.
With Python I don't get the same feeling, they seem more resistant to change in general. I don't think it's just that Python has more so-called "non-programmers" (meaning, data science and math people). Most of the libraries for these things are made by full-time programmers anyway. Perhaps the idea that "there's only one way to do it" brought about a "the current way is good enough, why rewrite in a new one?" attitude.
But anyway, I'm not familiar enough, just an outside perspective.
That was when I officially stopped recommending Python as a new language. It's just got too many things to explain now. And too many of those things to explain are like decorators, which requires two explanations, one to explain what it looks like they are doing when you use a decorator prepared by someone else, and a second explanation about what they are actually doing if someone ever wants to write one of their own.
It looks like "x.y" resolves the "y" value out of the instance, but actually, it's this complicated thing so that properties and other things work. It looks like the for loop simply ranges over the list, but it's actually this thing with iterators. It looks like calling this function "yields" a list but it's actually this magical thing where using a "yield" in the function body completely and non-locally changes the entire function. It looks like....
When Python was new and the list of this "it looks like..." was short, it was kinda need to be able to override __getitem__ for certain useful effects and such, but stacking on more and more of these with every release has created a real monster of comprehension.
Now it feels like Python has raised "it looks like... but actually..." to a design principle.
It looks like the match operator simply takes the argument you give it and sees if it is the same "shape" as the case statements, but it actually consults a new magic class variable and has this logic and that logic and....
This seems a little unfair. Testing frameworks are notorious for pushing the envelope of language constructs and weird introspection tricks. And if in this case you picked an exceptionally ripe example I'm not sure that should reflect on the entire language.
Back in the day people used to use the baroque nature of some frameworks such as Zope to bash Python but Zope was an outlier. There will always be some gorgonzola in any ecosystem.
I'm not sure i agree entirely: regardless of whether we're talking about decorators/annotations in Python/Java/.NET/... in the context of testing/web/... frameworks, any and every of those combinations have only caused me pain and headaches.
While i do agree in principle that it's nice to abstract away code with something that annotates a method and gives it additional meaning, in practice this sort of indirection usually has something to do with dynamic code calls, reflection or whatever the idiom is called in any particular language. All of those feel harder to debug and more brittle than just writing 3 lines of code instead of 1 annotation/decorator.
The problem with that, is that it's usually:
- hard to make your own decorators/annotations (especially if you want to extend their functionality)
- hard to introduce state, dependencies and just get them working nicely with the rest of your codebase and IoC/testing frameworks, or pass data between the logic behind them (from one decorator/annotation to another)
- hard to use them, as opposed to your "regular mode" of writing code, IDEs won't offer many refactoring options
- hard to debug them, especially in the case of wanting to examine their internals and use breakpoints
To me, it just generally seems like someone is trying to tack on functionality in ways that are hard to utilize, as opposed to just writing code in the first place, the very same thing that you see happen with generics, deeply nested inheritance and other "tricks" which are supposed to make one's life easier but do the opposite in practice in many cases.Personally, i enjoy libraries like Typer (https://typer.tiangolo.com/) for Python and think that it has plenty of use cases, but i've been hurt far too many times by similar functionality in non-trivial cases in every single language that i've used that supports something like that, so at best i'm guarded about utilizing them.
Java rant: But perhaps that's just because Java is a major source of pain for me in that regard. To give you a concrete example: i'm migrating Spring to Spring boot and someone used Jersey instead of RESTEasy as their JAX-RS implementation and now i need to transpose hundreds of API endpoint definitions (for example @GET to @GetMapping with bunches of parameters). If it were just Java code, i could probably use some clever refactoring in it with the help of my IDE, but now i'd have to figure out where the annotations come from, create a stub to replace them, call the proper stuff from within them and hope it works, unless the reflection that's used by them breaks. Not only that, but debugging is kind of hard when the source code and the actual logic behind said annotation is hidden below dozens of layers of Eldritch indirection and you have no chances of feasibly finding all of that stuff out.
More on topic, I do sort of feel like maybe the Python team should have made a Newthon (with a better name) or something, because Python is becoming a shambling beast. Starting as one of the world's most dynamic OO language and then shambling towards being a static language, maybe a bit of a functional language, it's crazy.
I often joke about "Katamari Dama-C++!", which has the irresistible phonetic allusion, but damn if Python isn't trying to keep up at this point.
If you don't like Python for some task... use something else suited for your task, then. It's not like using something else means you must put Python down and never use it again on pain of being shot or something. Stop waiting for Python to mutate into O'Caml or something and just use O'Caml or whatever.
Decorators though I totally agree; had some tooling at work that was gussied up to the point of impenetrability by gratuitous use of decorators. A colleague got so fed up with it he rewrote it in Perl, which was, shockingly, actually a maintainability improvement.
Same story for list comprehensions: it's almost beyond me how one can be against a simple single list comprehension. It's simple, readable, almost beautiful. Now if people start nesting 5 comprehensions it's not so nice anymore; but then is the language to blame (as in, would we really want to forbid nested comprehensions?), or the people?
Comprehensions are a little less magical - if anything, they are more explicit. If I want to create a list where each element is generated by some function over another list, I just say it. Doing so with loops is obscuring what I wanted to say in the first place. The problem with comprehensions isn’t so much the comprehension, but the obtuse ways people can use them to eke out performance by avoiding explicit loops.
I’m all for things that are closer to what a programmer means, but less keen on features that entail obscuring details that may come back to haunt the programmer later (I see this most often with decorators).
"But someone will make a typo one day resulting in = instead of ==" argument is nonsense as that hasn't been an issue since forever in other languages. You can either require parentheses if you want to use the value (the way C compilers want it) or just make it := to begin with.
I was once an avid Python user, but I switched to mostly Rust a few years ago and now write very little Python. A while ago I wrote some not quite trivial script-like code in Python, but ended up converting it to Rust before it finished running. The attraction of Python used to be that it was simple, but I now value robustness (through the type system in particular) much more than simplicity of the language.
Oh well, maybe in a few years someone will put in a PEP to use walrus := with match/case to give match assignment and give the old Python guard something to _really_ clutch pearls about.
x = match expr:
case 'a':
1
case 'b':
2
case _:
0
I think mainly it doesn't work very well in Python due to indentation problems. Kind of like how lambda is restricted to single expressions, instead of blocks. Here's what they say in the rationale PEP (https://www.python.org/dev/peps/pep-0635/):> Statement vs. Expression. Some suggestions centered around the idea of making match an expression rather than a statement. However, this would fit poorly with Python's statement-oriented nature and lead to unusually long and complex expressions and the need to invent new syntactic constructs or break well established syntactic rules. An obvious consequence of match as an expression would be that case clauses could no longer have arbitrary blocks of code attached, but only a single expression. Overall, the strong limitations could in no way offset the slight simplification in some special use cases.
You can dislike this language characteristic, but it's the right call to stay congruent.
This would make python both a viable choice for rapid prototyping, as well as easy to convert to much higher performance code (just by handling some typing conflicts), and potentially making python a mega-language used even more pervasively.
https://github.com/adsharma/py2many/blob/main/doc/langspec.m...
The tool can convert your code to C++/Rust/Go and give you a native binary which like you observe will run faster without the python runtime.
There are many open tasks that could use help.
Py2many is a transpiler, which converts source code of one language to source code of another language (which then needs to be compiled).
So far, Py2many seems to support Rust and C++14, and preliminary support for Julia, Kotlin, Nim, Go and Dart.
Py2many tries to primarily build a standalone C program without the python runtime, although it supports a --extension flag, where it generates a PyO3 extension for rust.
If you use additional type markings it can generate more performant C code.
BTW: I remember a blog article of someone who was using Cython as the primary language instead of Python.
As the sibling comment notes, transpiler is not a very useful term.
A thought assembler was a "compiler", where a "compiler" is a special type of transpiler that is only capable of producing machine code.
There are also many, like me, who find it to be a distinction without a difference. We call them all compilers.
One example is the TypeScript team:
> Let’s get acquainted with our new friend tsc, the TypeScript compiler.
https://www.typescriptlang.org/docs/handbook/2/basic-types.h...
Another is Bjarne Stroustrup. Back in the glory days of BIX, I posted a comment in one of the forums that said, "C++? That's a preprocessor, isn't it?" (At the time, the only C++ implementation was Cfront, which translated C++ to C.) Bjarne replied in no uncertain terms that Cfront was a compiler.
Longer version of the story with some discussion of other words and phrases:
https://news.ycombinator.com/item?id=15154994
Of course, the word "transpiler" had not yet been coined, but I have a feeling Bjarne still prefers to call Cfront a compiler, as the Wikipedia page does:
> Cfront was the original compiler for C++ (then known as "C with Classes") from around 1983, which converted C++ to C; developed by Bjarne Stroustrup at AT&T Bell Labs.
https://en.wikipedia.org/wiki/Cfront
So to my mind, a compiler doesn't have to generate a "compiled executable". It also doesn't have to take "source code" as input.
Most implementations of JavaScript and Java have at least two compilers: one that translates source code into bytecode, and a Just In Time Compiler that never sees the source code, but analyzes and profiles the bytecode while it runs, and translates sections of the code into machine language as it finds hotspots.
Now if you like the word "transpiler", as many do, I won't try to convince you otherwise. I just wanted to explain why some don't find it a useful distinction and prefer "compiler" for all these cases. It's a program that translates computer code from one language to another, whatever the form of those languages may be.
I suppose by my own argument I should also call an assembler a compiler! But no one does that... :-)
Friend fyi... the word "transpiler" to make a distinction of converting the source to another higher-level language rather than a lower-level machine language appeared as early as 1960s. In the 1964 paper, see the last page and the 2nd-to-last paragraph:
http://comjnl.oxfordjournals.org/content/7/1/28.full.pdf+htm...
>There are also many, like me, who find it to be a distinction without a difference. We call them all compilers. [...] Now if you like the word "transpiler", as many do, I won't try to convince you otherwise. I just wanted to explain why some don't find it a useful distinction and prefer "compiler" for all these cases.
I understand your point of why "compilers" already encompasses transpilers so the word "transpilers" seems redundant. But I'll try to explain why "transpilers" still endures. If you're not familiar with concept of "lumpers vs splitters", take a look at: https://en.wikipedia.org/wiki/Lumpers_and_splitters
Imagine if we did not have the word "transpiler". Language usage would still evolve to make a distinction via extra prefixes or suffixes or extra adjectives. Examples of alternative history might be:
- "compiler-to-asm" or CTA acronym, or "compiler-to-executable", or "native compiler"
- "compiler-to-another-high-level-language" or CTAHLA acronym, or "programming language translator"
But humans don't like to say or write verbose multi-syllable terminology and they'd eventually substitute a simpler word as shorthand for the distinction. That shorthand word might be "transpilers" or another word to serve the same purpose. And that's what happened in 1964 when AF Parker-Rhodes suggested the word "transpiler".
Similar concept of labeling Linux ext4, Microsoft NTFS, Apple HFS as "file systems" instead of just "databases" -- even though they are all conceptually "databases" which have same computer science concepts of keys, blocks, btree indexes. If we tried to police the language and insist that "file system" adds no useful distinction because it's a "database"... human language usage would still evolve with extra adjectives such as "operating-system-built-in-database-for-documents-and-arbitrary-files-etc" ... which is a cumbersome mouthful which motivates a shorthand terminology... perhaps call it "file system".
The use case for mypy is strictly for static type checking AFAIK.
Now, my dream would be for Python to actually enforce type annotations at runtime when they exist. Maybe add some kind of syntax to avoid doing it when it's expensive, e.g. `list[nocheck int]`, but for the love of god, enforce it, otherwise I feel like I'm just peppering my code with limp suggestions that only stand if some third party package can find all paths that lead to them, at which point I'd rather use a true statically typed language where the whole system doesn't look and smell like a janky afterthought.
Finally, CPython is very old. That means it made decisions around multithreading it still has to live with to this day whereas the JavaScript implementations have always been single threaded and added multithreading much later (learning from the experiences of other languages making the transition).
Dynamic typing is part of the problem but you’re right that it’s not the whole thing.
The bigger thing is that the Python runtime has a GIL because that's kind of how you solved the problem when multiple cores started becoming mainstream in the 90s. The Linux kernel had a similar GIL. It's an easy way to technically support running in a multi-threaded environment while reducing maintenance issues at the cost that you can never use more than 1 core at a time. JS runtimes don't have this problem because the runtime support layer is almost non-existent & there's no multi-threading support at all (indeed, V8 by itself comes with almost no JS APIs - there's a clear delineation between language runtime & "app runtime"). Service workers is extremely thin & works on channels instead of shared memory, further simplifying the design. The JS GC is typically written to support concurrent multi-threaded operation to minimize the "stop the world" phase runtime whereas Python's cycle detector needs to stop the world due to the GIL. These are all small decisions that add up.
https://www.typescriptlang.org/docs/handbook/2/narrowing.htm...
function isFish(pet: Fish | Bird): pet is Fish {
return (pet as Fish).swim !== undefined;
}
Which I've found extremely usefulProperly statically typed languages do not need run-time checks because the compiled code is already proven to be valid by the type-checker.
Okay, so there are some, though flow is pretty obscure, and flow-runtime (which is the part which does this AFAICT) is even more obscure. I don't think this is a popular feature.
> Properly statically typed languages do not need run-time checks because the compiled code is already proven to be valid by the type-checker.
If you want that, just enable strict mode for mypy.
And if you want to interface with code that is not stirctly typed, just use runtime type checks, which will inform the static type system:
foo: Any = untyped_code()
assert isinstance(foo, MyClass) # after this mypy will treat foo as instance of MyClass.
Same can be done with other types of checks also.So if you want runtime type checking that informs the type system, you have it. I certainly don't want runtime checks for what the type system has already checked for me.
But I can't do that if I want to do gradual typing. If I type a function, but the type system cannot figure out what's going to call it, that's when I want it to perform a check: at the boundary.
> I certainly don't want runtime checks for what the type system has already checked for me.
The type system can remove the checks if it can prove that they always pass. My point is that if it cannot prove that they will pass, I want them at runtime.
Unless you start using # type: ignore - the only way to interface with untyped code will be to use asserts or other checks (if + throw), i.e. perform checks at the boundary.
Well, what's a properly compiled language these days? Golang e.g. uses run-time dispatching due to the way interfaces work and, of course, garbage collection. Even C++ needs to do run-time dispatching of virtual methods. So the only languages left these days that do not use run-time checks are probably C and Fortran.
For example, in JS you might put these all over your code-base:
if (typeof x !== 'string') {
throw Error('Expected a string');
}I e said it before and I’ll say it again: runtime type checking saves you from the case where you pass None instead of an actual object but not much else. How wrong would your program have to be to pass an instance of Animal instead of an instance of BankAccount? And how quickly would it error out with “giraffe does not have an attribute routing_number”?
Does any mainstream language do this? For sure most don't, even statically-typed ones. I'm all for enforcement of types, but why does it need to happen at runtime?
I'd argue a good JIT compiler will always beat a static compiler, especially for a dynamic language using a complicated syntax such as python.
I think it's indicative that the nicest use is in AST processing. As someone that works on the occasional parser-y side project it looks pretty cool, but also, 97% of professional Python development is in data-science/application development. I suspect this is a case of the core language devs having a bias towards features that are useful to them as opposed to the community as a whole.
Apps designed pre ubiqutous pattern matching say little about whether pattern matching has good uses or not. More likely features alluding to pattern matching were culled early in the design process.
That said, I think your overall point is interesting. Will this new feature in Python mean that Python is used for certain domains it wasn't previously? I have my doubts, partly because it's a fairly slow byte-code interpreted language, and that limits its usefulness (I've seen a number of compilers written in OCaml, for example, which has pattern matching, but is also fast).
Python has to cater to a vast crowd, and geographers don't have the same needs as data scientists, web devs, 3D artists or sysadmin.
So it's logical some tools will be useful in some cases, and not some others.
Amusingly yesterday I was doing something in JS, and though, "ah, I wish I has this new PM feature from python there".
The day before, I had to deal with pedantic exception, and wish I had a pattern matching to make a case depending of the number of errors and their message. I had to do a hacky loop for something that had to do with structure, not repetition.
Just like the walrus, I don't think it's something that will pop up everywhere, all the time. But when it is useful, it will make the code nicer.
No. There is not a single field with " 97% of professional Python development". As a freelancer, I see python used for everything, everywhere. From education, to math, to biology, to geography, to web dev, to automation, to testing, to desktop app...
Powerful? I guess so.
Complicated? Definitely.
Confusing? You bet.
Good addition? Uhh, I really hope so.
The value of the concept really comes out when you start designing your types to be case...match-ed against. I am not the best person to explain this, but the value I mean is allowing more 'algebraic datatypes' to be used in python. This way you can have the structure of your data(types) more closely resemble your business logic. Making code more readable, and also making it harder to represent invalid states, which reduces the need to validate data. (I recall an article called "parse don't validate" the same idea applies here)
I think it will take a while for python users to start using this paradigm, but I expect it will be very useful for folks who know Rust, Haskell, or similar languages.
This is the problem with the recently added features in Python. As the author of the blog post correctly notes, the features are unintuitive and full of special cases.
You always have to know and recreate the steps that the unpacking machinery takes, there is nothing logical and declarative like in the ML family languages.
Pattern matching in SML or OCaml is simple and obvious.
These features are added to Python in order to give an impression that some "development" is happening. With the benefit that they prevent implementations like PyPy from catching up. Meanwhile a non-existing vaporware JIT from Microsoft is pushed as the future.
I would like to have seen some benchmarks for it, though, specially for things like the gaming loop, where, compared to the if-variant, the type of the event has to be checked for every case, since grouping is not possible.
This specific game loop is not adequate, since input events are rare, but if this were to be used with Scapy for example, it would really start to matter, like when iterating over raw socket events.
Yeah, the match is slower than if-else in that case because they're not grouped. However, the PEP points out that faster implementations are possible -- it would not have to check the type each time if it already knows it can't match. I guess some of those optimizations are left for later.
One other limitation from the perspective of people looking to transpile static python to other languages: match is a statement, not an expression in the language.
I've been looking to define an extension for transpilation purposes that doesn't break match in 3.10 but provides an avenue to transpile code to rust and other FP languages.
"In order to get all of the details about the tuple from the function one must analyse the bytecode of the function. This is because the first bytecode in the function literally translates into the tuple argument being unpacked. Assuming the tuple parameter is named .1 and is expected to unpack to variables spam and monty (meaning it is the tuple (spam, monty)), the first bytecode in the function will be for the statement spam, monty = .1. This means that to know all of the details of the tuple parameter one must look at the initial bytecode of the function to detect tuple unpacking for parameters formatted as \.\d+ and deduce any and all information about the expected argument. Bytecode analysis is how the inspect.getargspec function is able to provide information on tuple parameters. This is not easy to do and is burdensome on introspection tools as they must know how Python bytecode works (an otherwise unneeded burden as all other types of parameters do not require knowledge of Python bytecode).
The difficulty of analysing bytecode not withstanding, there is another issue with the dependency on using Python bytecode. IronPython [3] does not use Python's bytecode. Because it is based on the .NET framework it instead stores MSIL [4] in func_code.co_code attribute of the function. This fact prevents the inspect.getargspec function from working when run under IronPython. It is unknown whether other Python implementations are affected but is reasonable to assume if the implementation is not just a re-implementation of the Python virtual machine."
I'm not very convinced without further information - it sounds like it was already solved in CPython, and bug in IronPython? The whole PEP reads as if it's looking for excuses for removal.
I find this especially annoying for lambdas.
There are so many natural real-world applications to tagged union structures that I bet the ratio will look very different in a couple of years.
interface Foo { type: "foo"; name: string; }
interface Bar { type: "bar"; size: number; }
type Variant = Foo | Bar;
function tell(x: Variant) {
switch (x.type) {
case "foo":
// x now has type Foo
console.log(`The name is {x.name}`);
break;
case "bar":
// x now has type Bar
console.log(`The size is {x.size.toFixed(2)}`);
break;
default:
// statically unreachable
}
}
Hope to see this in Python type checkers as well. from typing import Union, NoReturn
class Foo:
name: str
class Bar:
size: int
Variant = Union[Foo, Bar] # or `Foo | Bar` in Python 3.10
def assert_never(x: NoReturn) -> NoReturn:
# runtime error, should not happen
raise Exception(f'Unhandled value: {x}')
def tell(x: Variant):
if isinstance(x, Foo):
print(f'name is {x.name}')
elif isinstance(x, Bar):
print(f'name is {x.size}')
else:
assert_never(x)
mypy returns an error if you don't handle all cases.It will be much more useful as an alternative to dispatch dictionaries, which are more a side-effect of the language lacking any case based control-flow.
Though considered from that perspective, it’s actually a bit of a disadvantage that the new match statement is a part of the language. It’s nice to dynamically extend (either as a first party developer or third party user) that sort of dictionary, effectively adding new branches to the implicit switch statement. A match statement is going to be fixed as written. Probably a readability and consistency win to remove spooky-action-at-a-distance, but at the cost of a bit of (ugly, pragmatic) usability
But if you're not matching structure I really don't see what's wrong with plain old if...elif? And you can easily add more complex tests like "in" clauses this way too:
cmd = args.command
if cmd == 'push':
# do push
elif cmd in ['pull', 'pul']:
# do pull
else:
error(f'command {cmd!r} not yet implemented')I didn't RTFM but I did CFTFM (ctrl-F'd the FM).
For the enum switch, match generates very similar bytecode (use dis.dis to see) and the execution time is almost identical.
For the structural matching, match is significantly slower, which surprised me a bit (it's almost twice the number of bytecode instructions). That said, we're down in the nanoseconds range -- use of this feature shouldn't be about ns-level performance, but code clarity. I suspect "match" will get faster over time as they add optimizations or specialized bytecode for it.
Thanks for resubmitting, chmaynard!
Consider their first example, interpreting a command in a textual game. In many places the logic is divided into a few pieces: check that there are the right number of arguments, do something with some of them, pull off the rest, etc. It's really easy to make mistakes with this kind of code, where your conditions don't quite align. With pattern matching, you write a single pattern for a case, and I would expect this to give many fewer bugs.
I find that 99% of the time that you need to use threads for CPU work (not just IO), that means that you care enough about performance that your Python just calls into some C library (C extension libraries e.g. numpy, scipy, or even bits of the standard library e.g. zipfile operating on an in-memory buffer). The thing is, all these calls (almost) all release the GIL anyway. So threading support in Python is already meaningful, unless you want to run pure-Python CPU intensive code on multiple threads, which I'll admit does happen but in my experience is rare.
Unfortunately we're mostly dealing with complex stuff happening during graph traversals. There is no "inner loop" for us -- just the "outer loop" of the graph traversal. So the only way to get our code to run in parallel is to port the whole algorithm to C. But that's 150k+ lines of codes, there will be barely any Python remaining :(
IMHO Python is a trap, it's easy to get started with but you can stuck in a position you can't get out of without a full port to a different language. (and yes, we've already tried multi-processing and using shared-memory segments in the C code portions to avoid re-doing work in every process, ... but after more than a man-year of work, >80% of our application remains single-threaded thanks to the GIL, even though in any other language it would be trivial to parallelize the independent graph traversals)
It's neat to have a scripting language for fast prototyping but we would be much much better off if we had picked Javascript instead of Python (the ability to have multiple interpreters in the same process running in parallel would solve our issues).
As I said, my experience is that this is not normally a problem in practice. Clearly that's not your experience! I'm not sure whether I've been unusually lucky or you've been unusually unlucky.
The structured nature and reified syntax works alongside the ability to quickly change the spec, as otherwise you would need to pick apart and reverse-engineer someone's in-the-moment if/elif (likely nested) when there's a change. The structure being based on first-class concepts in the language means that its meaning is common to all code, whereas the post went out of its way to specify how you'd "organise this differently as if/elif", which is clearly something requiring skill and experience (which experts always overestimate the level of), and would produce varying results depending on the person.
This will be taught at a more intermediate level of Python, leaving the complexity only skin-deep, akin to teaching pattern matching in Haskell or such. After all, the full complexity isn't necessary for most users, the vast majority of people will only read the tutorial PEP and not the full description and grammar, while those advanced users can make use of this to make life even easier. Combined with typing, this means tools can do more exhaustive checks and catch bugs pre-emptively,
In my opinion, this will be a boon for clean programs that are easy to read and pleasant to write. In line with what the post says about usage settling down in a couple years, libraries and APIs will find the right balance with time, whilst making it easier and faster to write the thing you meant correctly, especially as this is a boon to tooling since you're quite literally indicating the structural intention of your code.
I see this alongside decorators (used sparingly), comprehension (used one or two levels deep), and types as helping people write the code they want to write, what they mean to write, and the language helping them in that endeavour. Just as we finally have dict.__or__ and dict.__ior__, it's something that lets people write what they intend, rather than the scaffolding that is needed to support it. I personally think that's a positive, and the kind of feature that will become very natural for many people to use in Python even if they don't have a background in languages that have pattern matching.
I'm at the point where I'm willing to pay someone to write a preprocessor that'll let me put switch/case statements in my code and then run the preprocessor and have it convert them to if/elif/else statements.
Stop trying to look more important and smart by using more (unnecessary) words.