Proposing a struct syntax for Python
snarky.ca
snarky.ca
https://docs.python.org/3/library/dataclasses.html#module-da...
..and the syntax is way better, handling more members, with room to document, handling methods, and not conflicting with Mojo or stdlib
FWIW I think this article mentioning them at all is sort of weird because the article doesn't propose them either, lmao
class A(BaseModel):
status: Literal["success"]
class B(BaseModel):
status: Literal["error"]
class C(BaseModel, ABC):
__root__: Union[A, B] = Field(discriminator="status")
aorb = C.parse_obj(blah…).__root__
if aorb.status == "error":
# aorb will type check as B.I was looking for a solution in pydantic that is wildly relevant. Thanks, I didn’t realize that was implemented/existed.
My only objection is using ABC in your MRO. Haha, just a little bit overloaded with abstract base classes and using A, B, and C.
I got it eventually! Cool.
I don't know if type checkers can figure out if a pattern match statement is exhaustive yet though, that would be interesting.
[1]: https://stackoverflow.com/questions/16258553/how-can-i-defin...
Then again, you can just create newtype wrappers, and that's essentially the same thing as different constructors for a sum type.
To that end, I'm curious what defaults you didn't use?
> - Typing is optional (there are still folks who don't want to lean into that and it simply isn't always necessary)
> - Class construction should be faster (which intuitively you would think shouldn't matter, but I have talked to folks where this is an actual concern, especially when startup time is critical)
> - Syntax typically leads to better tooling support since there's no ambiguity
> - Easier to teach than (data)classes, so can act as a stepping stone towards classes
> - Better semantics than dataclasses have by default (at least in my opinion )
Honestly why would you cater to people who want to write worse code? Should Python have a mode with implicit type coercion too because some people can't be bothered to type `int()`?
> Class construction should be faster.
It's Python. You accepted your code would be extremely slow going in. Fiddling with the syntax isn't going to make it fast.
Not that it's not worth spending some effort to optimize, of course. But the lower bound is quite high even if you're very careful.
Someone who really wants to avoid type annotations at all costs can just use ": object" everywhere
Who are people who are against typing in this day and age?
People who don't find the benefits of typing are worth the overhead?
Python is, after all, a dynamically-typed language, so it is, surely, not too surprising that those who use it might include those who want a ... dynamic language ?
If it's too much to take the joys of Java are freely available.
What overhead, specifically?
I've found it reduces overhead. Rather than using English to describe the types, in the comments/docstrings, I can just type the terse type hint in the function definition, and leave the rest to the documentation renderer. If you don't document anyways, sure.
But not having methods would produce APIs that feel really non-idiomatic. Essentially all Python types have methods, even immutable 'pure data' types. str has methods (though some things like `len` and regex search, which are methods in other languages, are functions in Python). int has methods (though they're rarely used). Immutable standard library classes – say, datetime or Path – have lots and lots of methods.
And inheritance… well, I could live without it but I've made good use of inheritance with dataclasses.
And as others have mentioned, the syntax is awkward.
I think my ideal solution would basically be dataclasses, but implemented in C and with more syntax sugar:
# Instead of "@dataclass\nclass Foo:"
struct Foo:
# Allow untyped fields:
x
# Make this do what you'd expect, instead of
# requiring "field(default_factory=list)":
y = []
# Of course, types are still allowed:
z: int = 0
Methods and inheritance would be allowed as usual.It might be interesting if these structs acted as true value types, so that e.g. just assigning an instance from one variable to another would make a copy of all the fields. That way you could mutate the object without worrying about sharing. But arguably, having both value types and reference types would overcomplicate the mental model.
- it's immutable
- doesn't have methods so no overloaded operator or dynamic attribute/index accesss
- shape is known and is set in stone
That's why Mojo has them, and why they are the bread and butter in Rust.
It's one of the step in the direction of a faster Python.
The author is proposing a worse dataclass, because they don't want to ignore positional indexing on namedtuple?
Say I want to model in Python some collection of key-value pairs. It's a singleton -- I don't need a class or namedtuple that can be instantiated. So for example, I just want something like:
SingletonDataContainer Config:
root_directory: str = "/tmp"
should_frobnicate: bool = False
So of course, it could be a dict. But I want type annotations, and dotted attribute access syntax (and pretty syntax-highlighting!). It's tempting to use a class with class attributes for this, but that seems wrong since then you end up with something that can be instantiated, but which you never actually want to instantiate.A module variable set to a SimpleNamespace in the typing package is also an option.
from importlib.machinery import ModuleSpec
import importlib.util
import sys
spec = ModuleSpec("mymodule", loader=None)
module = importlib.util.module_from_spec(spec)
module.__dict__.update({"a": 42})
sys.modules["mymodule"] = module
import mymodule
print(mymodule.a) # -> 42
That's if you want a proper module. But modules can be any type and it's officially supported! import sys
from types import SimpleNamespace
sys.modules["mymodule"] = SimpleNamespace(b=37)
import mymodule
print(mymodule.b) # -> 37
# ^^^ This is literally just calling __getattr__ on the module object. Do with that what you will.
This will cross files as well but you need to make sure the code to create the module runs before you try importing it! glhf> Python? Rules? Unheard of.
OK, since it's Christmas, I'd like to be able to write
Module MyModule
a = 42
b = 37 import sys
class DataModule:
def __init_subclass__(cls):
sys.modules[cls.__name__.lower()] = cls
class MyModle(DataModule):
a = 42
b = 37
import mymodule
mymodule.a == 42
Or as a class decorator def datamodule(cls):
sys.modules[cls.__name.__.lower()] = cls
@datamodule
class MyModule:
a = 42
import mymodule
mymodule.a = 42 class Config:
root_directory: str = "/tmp"
should_frobnicate: bool = False
if __name__=="__main__":
print(f"{Config.root_directory=}\n{Config.should_frobnicate=}") const config = {
rootDirectory: "/tmp",
shouldFrobnicate: false,
}
I'm far from a python expert, more like a TS dev who occasionally tries to write Python, but when I'm in that situation I find that writing one-off classes seems to be python's way of doing things. I'm not a super big fan of it myself, but it works: from dataclasses import dataclass
def _create_config():
@dataclass(kw_only=True)
class Config:
root_directory: str = "/tmp"
should_frobnicate: bool = False
return Config()
config = _create_config()
I experimented with making a decorator that could do something like this, but never got it working with `mypy` so I gave up: @iife
def config():
@dataclass(kw_only=True)
class Config:
root_directory: str = "/tmp"
should_frobnicate: bool = False
return Config()
the `@iife` would make it so that `config` was a variable, not a `def`, but yeah the whole thing stank so I gave upIt should be possible to get simple attribute access while not creating a new nominal type and not going outside the Python stdlib.
I'm constantly using NamedTuples when what i really want is an anonymous struct.