Approximating sum types in Python with Pydantic
blog.yossarian.net
blog.yossarian.net
That is, you can approximate Rusts's enum (sum type) with pure Python using whatever combination of Literal, Enum, Union and dataclasses. For example (more here[1]):
@dataclass
class Foo: ...
@dataclass
class Bar: ...
Frobulated = Foo | Bar
Pydantic adds de/ser, but if you're not doing that then you can get very far without it. (And even if you are, there are lighter-weight options that play with dataclasses like cattrs, pyserde, dataclasses-json).[1] https://threeofwands.com/algebraic-data-types-in-python/
Also I’d add msgspec to your list at the end. Lightweight and fast, handles validation during decoding.
(Python’s annotated types are very powerful, and you can do this and more with them if you don’t immediately need ser/de! But they also have limitations, e.g. I believe Union wasn’t allowed in isinstance checks or matching until a recent version.)
It also supports complex structures like union types, lists, etc. I used it to create cresset [2], a package that allows building Pytorch models directly from config files.
[1]: https://pypi.org/project/confactory/ [2]: https://pypi.org/project/cresset/
It was also an order of magnitude slower than other libraries, and at the time all these libraries were much slower.
@enumclass
class MyEnum:
class UnitLikeVariant(Variant0Arg): ...
class TupleLikeVariant(Variant2Arg[int, str]): ...
@dataclass
class StructLikeVariant:
foo: float
bar: int
# The following class variable is automatically generated:
#
# type = UnitLikeVariant | TupleLikeVariant | StructLikeVariant
where the `VariantXArg` classes are predefined.Do you have anything public that elaborates on this?
If you don’t care about types and just want ser/de that’s great, but I think it’s clearly on topic here to care about types.
Essentially, if this is a feature you must have, Python seems like the wrong language. Maybe if you only need it in spots this makes sense...
In the same way that SQLite bills itself as the better alternative to fopen(), Pydantic is the better alternative to json.loads()/json.dumps().
Every step that takes Python in that direction is a mistake, because if we need to make a huge commitment, Python probably isn't the right language. A large part of the appeal of Python is that it is easy to learn, easy to bring devs up to speed on if they don't know it, easy to debug and understand. That's why people use it despite its performance shortcomings, despite its concurrency issues, etc. (That and the benefit of a large and fairly high quality library.)
I think Python's journey is very similar to Go in this regard where as the language matures and more people start using it for large applications you start having to compromise on the ease of on-boarding in favor of the users who are trying to get work done. Both Python and Go added generics around the same time.
This like saying "instead of understanding Python, you have to understand a bunch about SQLAlchemy and ORMs" or "instead of understanding Python, you need to understand GRPC and data streaming."
Ultimately every library you add to a project is cognitive overhead. Major frameworks or tools like sqlalchemy, Flask/Django, Pandas, etc. have a lot of cognitive overhead. The engineering decision is whether that cognitive overhead is worth what the library provides.
The measurement of worth is really dependent on your use case. If your use for Python is data scientists doing iterative, interactive work in Jupyter notebooks, Pydantic is probably not worth it. If you're building a robust data pipeline or backend web app with high availability requirements but dealing with suspect data parsing, Pydantic might be worth it.
The phrasing of "The engineering decision" in your reply is telling -- you are coming from it as an engineer. But I'm looking at the population of Python programmers, which extends far beyond software engineers. The more such people have to learn, the more problematic the language becomes. Python succeeded despite not being a statically compiled language with clear typechecking because there is an audience for which those aren't the critical factors.
As I said in another response, it reminds me of what happened to Java. Maybe that's just my own quirk, but none of these changes are free.
What you have to know depends on where you're working and what you're doing. You don't have to know GRPC Python libraries, unless it's a company that uses GRPC for internal communication. You don't have to know Flask unless you're building a REST API using Flask. You don't have to know beautifulsoup unless you're building a web scraper. You don't have to know Pydantic unless you're working on a project that uses Pydantic for data validation.
> The phrasing of "The engineering decision" in your reply is telling -- you are coming from it as an engineer. But I'm looking at the population of Python programmers, which extends far beyond software engineers.
You don't have to be a software engineer to make an engineering decision. When a data scientist uses conda because they don't want to manage their Python environment manually, but runs into performance issues on production because their Docker containers are multiple gigabytes larger than they should be -- that's the result of an engineering decision. When a business analyst writes a Python script and manually installs packages without a requirements file, then tries to get it running on a new computer 8 months later but can't because they forget which package versions they used -- that's the result of an engineering decision. So when you deploy your code without any data validation that runs fine now, but breaks in unexpected ways next week because the result of an external REST API you're calling changed unexpectedly...
> The more such people have to learn, the more problematic the language becomes. Python succeeded despite not being a statically compiled language with clear typechecking because there is an audience for which those aren't the critical factors. ... none of these changes are free.
I agree with all this, which is why I said that the engineering decision is deciding whether or not the cost is worth it. Different projects, companies, and people will have different needs.
Your original assertion was that "if this is a feature you must have, Python seems like the wrong language" -- but this contradicts what you're saying. The overhead of learning a single Python library is far, far less than, say, introducing Rust into a company that only uses Python for everything else.
I think my assertion, less pithily, was "if having the best type-checking system was critical to you, probably you wouldn't pick Python". And I think that's correct. People pick it for other features.
I didn't say I hated having the option. I expressed reservations, which I still have.
Agree, but the only situation where as a developer you can pick a language/ecosystem on its own merits, independently of anything else, is on personal projects. Even if you're a startup CTO building a greenfield app, you have account for hiring and train developers. It's perfectly sensible that you would want to use Python + mypy/pyright/pydantic/etc for extra robustness since it's easy to find Python devs, with a relatively small learning curve if they haven't use those tools, vs going full-on Rust or Haskell, which would require much more rare + expensive people and/or a much longer training period.
Also, here it is claimed the library should be part of the language, and at the same time it us assumed it is too complicated for the users to understand. It seems like the feature being a library solves this, if we let go of the self-imposed requirement of it being part of the language.
I don't claim Python developers cannot understand it. But every additional thing adds to the cognitive burden.
Also, even within the constrains of C++98 type system, expressivness wasn't something C++ was lacking.
I really like some of the languages you mention, but most of them were not popular, especially in Python domain at the time (scripts and web servers).
The main contenders against Python then were Perl (timtowtdi vs Python one way) and Java.
Java and C++ were strongly typed, but lacking most of the nice things of Haskell etc (at least Java). C++ is very expressive (prob more because of templates than the type system, but I will happily concede this one).
For scripts and web backends, C++ and Haskel/ml were not popular. This leaves Java, perl, php and similar, and at the time neither had advanced type systems in the ergonomic way that it is now expected.
There are a lot of people out there writing Python, and a lot of them identify as analysts, (non-software) engineers, scientists, and so on. Not developers. Some of them write immaculate code. Some of them don’t know about git, functions, or commandline arguments, so their code is one long script with big chunks they comment or uncomment depending on what they’re trying to do. The latter are a big constituency for me. Plain type annotations are great in this context, because they place no burden on the user at all. All they have to do is ignore them. Best case, they notice that function f returns a list of floats rather than an ndarray and that saves me having to explain what’s going on.
IMO a library that provides regular functions and values that follow the rules of the language adds zero cognitive overhead. Frameworks that change/break the rules, that let you do things that you can't normally do with regular values, or don't let you do things that you normally could do, are the ones that add overhead, and it sounds like Pydantic is more in that category.
In practice, however, Pydantic is one of the most popular packages/frameworks in Python because people do in fact need this kind of complexity. In particular, it makes wrangling complicated object hierarchies that come from REST APIs much easier/error prone.
> Essentially, if this is a feature you must have, Python seems like the wrong language.
While I don't disgree in the absolute sense, there are constraints. You can't just switch language or change the problem you're solving. If you have the need for more type safety, then this is a price worth paying.
There are some people resisting type checks in python but I think fewer and fewer. I dont think people refusing to learn basic concepts and libraries are a reason to not use something.
Also, I am not a big fan of not doing something useful because we need to do a bit of learning. It seems like a variant of "we have always done it this way". Plus it is a strawman attributed to python developers, IMO.
type committer =
InnerCircle of string
| NPC of string
| Dissenter of string
;;
type coc_reaction =
DoNothing
| ThreeMonthsWithoutHumiliation
| PublicDefamation
;;
let adjudicate = function
InnerCircle _ -> DoNothing
| NPC _ -> ThreeMonthsWithoutHumiliation
| Dissenter _ -> PublicDefamation
;;
# adjudicate (InnerCircle "Wouters");;
- : coc_reaction = DoNothing
# adjudicate (Dissenter "Peters");;
- : coc_reaction = PublicDefamation
Just use another language, also for social and professional reasons. @dataclass
class InnerCircle:
s: str
@dataclass
class NPC:
s: str
@dataclass
class Dissenter:
s: str
Committer = Union[InnerCircle, NPC, Dissenter]
class CocReaction(Enum):
DoNothing = auto()
ThreeMonthsWithoutHumiliation = auto()
PublicDefamation = auto()
def adjudicate(c: Committer) -> CocReaction:
match c:
case InnerCircle():
return CocReaction.DoNothing
case NPC():
return CocReaction.ThreeMonthsWithoutHumiliation
case Dissenter():
return CocReaction.PublicDefamation
Although in reality you'd likely model Committer as a product of a status and name, and adjudicate as a map of status to reaction, unless there are other strong reasons to make Committer a sum. from typing import Literal
class _FrobulatedBase:
kind: Literal['foo', 'bar']
value: str
class Foo(_FrobulatedBase):
kind: Literal['foo'] = 'foo'
foo_specific: int
class Bar(_FrobulatedBase):
kind: Literal['bar'] = 'bar'
bar_specific: bool
"kind" overrides symbol of same name in class "_FrobulatedBase"
Variable is mutable so its type is invariant
Override type "Literal['foo']" is not the same as base type "Literal['foo', 'bar']"
https://pyright-play.net/?code=GYJw9gtgBALgngBwJYDsDmUkQWEMo...The original example code with Enums doesn't type-check either, and for the same reason:
If the type checker allowed that, someone could take an object of type Foo, assign it to a variable of type _FrobulatedBase, then use that variable to modify the kind field to 'bar' and now you have an illegal Foo with kind 'bar'.
However, I think that's possibly a bug :-) -- I agree that narrowing a literal via subclassing is unsound. That's why the example in the blog used `str` for the superclass, not the closure of all `Literal` variants.
(I use this pattern pretty extensively in Python codebases that are typechecked with mypy, and I haven't run into many issues with mypy failing to understand the variant shapes -- the exception to this so far has been with `RootModel`, where mypy has needed Pydantic's mypy plugin[2] to understand the relationship between the "root" type and its underlying union. But it's possible that this is essentially unsound as well.)
[1]: https://mypy-play.net/?mypy=latest&python=3.12&gist=f35da62e...
def frotz(x: Frobulated) -> str:
return f”{x.value} is the value of x”Something I've wondered of late. I keep seeing these articles pop up and they're trying to recreate ADTs for Python in the manner of Rust. But there's a long history of ADTs in other languages. For instance we don't see threads on recreating Haskell's ADT structures in Python.
Is this an artifact of Rust is hype right now, especially on HN? As in the typical reader is more familiar with Rust than Haskell, and thus "I want to do what I'm used to in Rust in Python" is more likely to resonate than "I want to do what I'm used to in Haskell in Python"?
At the end of the day it doesn't *really* matter as the underlying construct being modeled is the same. It's the translation layer that I'm wondering about.
That's not to say it's bad, or a problem. If it gets more people into these concepts that's great.
I think so, in the sense that Rust has successfully translated ADTs and other PLT-laden concepts from SML/Haskell into syntax that a large base of engineers finds intuitive. Whether or not that’s hype is a value judgement, but that is the reason I picked it for the example snippet: I figured more people would “get” it with less explanation required :-)
Another way of phrasing my query is that given these are all basically ML-style constructs, why would the examples not be ML? And I was assuming the answer to that is "the sorts of people reading these blogs in 2024 are more familiar with Rust"
data Thing
= ThingA Int
| ThingB String Bool
| ThingC
To me, the above syntax takes away all the noise and just states what needs to be stated.My point being—you see articles about ADTs involving non-rust languages all the time. Why single rust out?
I would say familarity, and lack of exposure to programming languages in general.
As a stretch, I've seen Rust content where people claim that Rust has successfully popularized a handful of relatively obscure PLT concepts. But this is a much, much weaker claim than Rust innovating or inventing them outright, and it's one that's largely supported by the size of the Rust community versus the size of Haskell or even the largest ML variant communities.
(I say this as someone who wrote OCaml for a handful of years before I touched Rust.)
Here is another common one, "It would be great a Rust like but with GC".
What in this phrase suggests or implies that Rust has innovated something that an earlier FP language actually did? Something that resembles Go's managed runtime but with Rust's sum types seems like a very reasonable thing to want, and doesn't exist per se without buying either into a very foreign syntax and thus a much smaller community and library ecosystem.
(Or as another phrasing: what is actually wrong with someone saying this? Insufficient credit given to other languages? Do people apply this standard to C with BCPL and ALGOL? I haven't seen them do so.)
I don't think there's anything wrong per se. Although I do think it contributes to the sentiment that people may be ascribing things as being novel to Rust, even when not intended as in this case. To be fair, that's what sent me down the mental path earlier that prompted this subthread. And that's when I figured it was more a matter of being the implementation most likely to resonate with the audience.
And I don't think it's a matter of needing to give credit to other languages. But phrasing it like "Something with a managed runtime, but with sum types" is generic enough, unless there's something specific about either of those. For instance the phrasing I gave does exist in plenty of places, but perhaps "Something that resembles Go's managed runtime with sum types" perhaps does not. I don't know enough about Go to say that.
In other words, is there something specific about *Rust*'s sum types that one is after in this example? Or just the concept of sum types.
I think, concretely, it's the fact that Rust's syntax is more intuitive to the average engineer than ML or Haskell. Maybe that's a failure of SWE education! But generally speaking, it's easier to explain what Rust does to someone who has taken a year or two of Java, C, or C++ than to explain ML to them.
And that stopping point I think is where the perception of Rust's popularity on sites like HN is much higher than in the general public. And by that I mean people who at least grok, if not use, Rust and not people who like the idea of Rust.
For instance, keep in mind that even during the heyday of Scala here on HN the rest of the JVM world was complaining that Scala syntax was too arcane.
The contexts where it pops up, it is as if it would be yet to come, such language.
Speaking of C and BCPL, indeed we do, because many wrongly believe in this urban myth, that without them there was nothing else as high level systems programming languages, even though JOVIAL came to be in 1958, followed by ALGOL and PL dialects, Bootstrap CPL was never planned to be used beyond that purpose, and there was a rich research outside Bell Labs in systems programming in high level languages.
Instead we got stuck with something that 50 years later are still trying to fix, with Rust being part of the solution.
I don't understand why you think this: we explain things all the time without presuming that the particular choice of explanation implies ignorance of a preceding concept. In high school physics, for example, you wouldn't assume that your teacher doesn't know who Ptolemy is because they start with Newton.
The value of an explanation is in its effectiveness, not a pedantic lineage of the underlying concept. The latter is interesting, at least to me, but I'm not going to bore my readers by walking them through 65 years of language evolution just to get back to the same basic concept that they're able to intuit immediately from a ~6 line code snippet.
(It's also condescending to do so: there's no evidence whatsoever that Rust's creators, maintainers, community, etc. aren't familiar with the history of PL development.)
Some of us are tired of cleaning up after the inevitable messes these developers leave behind.
Please be a little bit more charitable with how you read comments. The core observation here is that “Rust is completely novel” is not actually something that Rust practitioners, including junior engineers, actually say. Nobody has said it in this thread, and nobody has even provided a single example of somebody saying it.
typedload does this without need to pass a "discriminator" parameter.
Just having the types with the same field defined as a literal of different things will suffice.
I've also implemented an algorithm to inspect the data and find out the type directly from the literal field, to avoid having to try multiple types when loading a union. Pydantic has also implemented the same strategy afterwards.
typedload is faster than pydantic to load tagged unions. It is written in pure python.
edit: Also, typedload just uses completely regular dataclasses or attrs. No need for all those different BaseModel, RootModel and understanding when to use them.
I agree with you that the article would have been improved if they'd used real-world examples, e.g. a ContactMethod type that has Address or PhoneNumber or something like that.
I've also played around with writing my own dataclass/data conversion library: https://github.com/hexane360/pane
https://github.com/adsharma/adt contains a small enhancement for @sealed decorator from the excellent upstream repo.
https://github.com/py2many/py2many/blob/main/tests/cases/sea... https://github.com/py2many/py2many/blob/main/tests/expected/...
More and more Java seems to be not that bad after all.
Data pipelines & ML workflows are already pseudo immutable and pseudo functional.
I very much credit Java for having turned a generation away from static typing, dynamic typing did get buoyed by a combination of moore’s law and good press but could never have done it without Java having smothered the other side and being dreck.
Go does have a poor type system, but it has nowhere near the verbosity of early aught Java: local type inference, free functions, any number of (public) types in the same file, closures, type definitions (terse and easy newtyping), iteration (if only for builtin types until 1.23), etc...
And that's without considering the cultural side of requiring two different implementations of every type (one interface and one impl) or XML-oriented programming for bindings "improved" by unreliable parsers going through method comments.
"The key point here is our programmers are Googlers, they’re not researchers. They’re typically, fairly young, fresh out of school, probably learned Java, maybe learned C or C++, probably learned Python. They’re not capable of understanding a brilliant language but we want to use them to build good software. So, the language that we give them has to be easy for them to understand and easy to adopt."
"It must be familiar, roughly C-like. Programmers working at Google are early in their careers and are most familiar with procedural languages, particularly from the C family. The need to get programmers productive quickly in a new language means that the language cannot be too radical."
We are all well aware of the blue collar goals of Java 1.0.
Now we are at Java 23 EA, where the distance to Go's type system is even greater.
We have Go's failure to learn from history of programming languages, ironically following Java's missteps on generics (reaching out to the same Haskell folks that helped with Pizza compiler), with warts of its own with magic string formats for timestamps, const iota dance for enumerations, magic types and tagged structs.
Had Rust become mature one or two years earlier, and most likely Docker and Kubernetes would have pivoted from Python and Java respectively into Rust instead, with Go's fate being the same as Limbo.
Had it not been the case, and it would have been as successful in the industry as Oberon-2 and Limbo, its two main influences, and (simplifying the actual historical facts) previous work from the authors.
Copilots have removed the effort overhead. Python's conciseness limits the syntactic overhead. Lastly, the emphasis on primitives and simple types limits cognitive overhead.
Static typing is winning precisely because it avoids the issues of 2010s-era Java.