- Faster attribute access: your code is faster
- Slotted classes take less RAM, less L1 cache pressure, your code is faster
- Wrist friendly to with .foobar instead of ["foobar"]
- Runtime error if you misspell an attribute name
- Faster attribute access: your code is faster
- Slotted classes take less RAM, less L1 cache pressure, your code is faster
- Wrist friendly to with .foobar instead of ["foobar"]
- Runtime error if you misspell an attribute name
For me it’s mostly about .attribute being more in line with the rest of the language. Kwargs aside, I find overuse of dicts to clunky in Python
https://wiki.python.org/moin/UsingSlots
Whether or not the performance matters...well that's somewhat subjective since Python has a fairly high performance floor which makes performance concerns a bit of a, "Why are you doing it in Python?" question rather than a, "How do I do this faster in Python?" most of the time. That said it _is_ more memory efficient and faster on attribute lookup.
https://medium.com/@stephenjayakar/a-quick-dive-into-pythons...
Anecdotally, I have used Slotted Objects to buy performance headroom before to delay/postpone a component rewrite.
And yes __slots__ improve perf, but it’s about avoiding the __dict__ access, which hits really generic hashing code and then memory probing more than it is about L1 cache
Where __slots__ are most useful (and IIRC what they were designed for) is when you have a lot of tiny objects and memory usage can shrink significantly as a result. That could be the difference between having to spill to disk or keeping the workload in memory. E.g., Openpyxl does with a spreadsheet model, where there could be tons of cell references floating around
> The __slots__ declaration allows us to explicitly declare data members, causes Python to reserve space for them in memory, and prevents the creation of __dict__ and __weakref__ attributes. It also prevents the creation of any variables that aren't declared in __slots__.
Emphasis:
> prevents the creation of __dict__ and __weakref__ attributes. It also prevents the creation of any variables that aren't declared in __slots__.
In short, if you create a slotted object with __slots__ it sends you down a fairly orthogonal object lifecycle path which does not create or use __dict__ in anyway. This obviously has drawbacks/limitations like not being able to add new members to the object like a normal Python object.
From the second article:
> However, if you have __slots__, the descriptor is cached (which contains an offset to directly access the PyObjectwithout doing dictionary lookup). In PyMember_GetOne, it uses the descriptor offset to jump directly where the pointer to the object is stored in memory. This will improve cache coherency slightly, as the pointers to objects are stored in 8 byte chunks right next to each other (I’m using a 64-bit version of Python 3.7.1). However, it’s still a PyObject pointer, which means that it could be stored anywhere in memory. Files: ceval.c, object.c, descrobject.c
Which I think addresses your concern about parent dict access...but I could also be misunderstanding your point.
depends on the type in question. If you are fetching and operating on a large number of records then it can matter. But otherwise the answer is more often that it does not really matter.
In many ways it matters more because it’s Python.
I’ve met a lot of teams throughout my career who struggle daily with a badly performing Python codebase. You can write a no-frills web service in c#, go, rust or JavaScript. And, so long as you don’t do anything stupid, it’s usually plenty fast enough from day 1 to handle your users. But in my experience, the same isn’t true of Python. I’m sure Python web services can be made to run ok, but because it’s slow by default, I bet a lot more time is spent optimising Python programs around the world than optimising JavaScript.
A brute force O(N) in C++ may be fast enough, in a situation where you need to use O(logN) to get the equivalent speed in Python. Squeezing out a few extra percent from a O(N) in Python by using slots will not be enough.
Of course that doesn’t mean you shouldn’t leave performance on the table if the optimizations have noticeable effects.
T = TypeVar("T")
U = TypeVar("U")
V = TypeVar("V")
P = ParamSpec("P")
def modelargs(model: Callable[P, U]):
def _modelargs(func: Callable[[T], V]) -> Callable[P, V]:
def __modelargs(*args: P.args, **kwargs: P.kwargs) -> V:
return func(model(*args, **kwargs)) # type: ignore
return __modelargs # type: ignore
return _modelargs
class MyModel(BaseModel):
foo: str
bar: int = 4
@modelargs(MyModel)
def test_func(model: MyModel):
print(model.foo, model.bar)
return 4
test_func(foo="Hello", bar=20) # -> prints Hello 20
If you look in your editor you'll see that the type signature for test_func is `(*, foo: str, bar: int = 4) -> int`. It's unfortunate that you have to write the model type twice but in exchange you don't have to write the args twice.Pydantic is targeting other use cases. The point of TypedDicts is compile-time safety without run-time overhead. Pydantic is useful for a lot of things, but performance isn't exactly its strong suit (written as of 2.9.2, I was just revisiting it earlier this week).
Anyway, in the same spirit of function signature hacking, I've found the following useful for "inheriting" them:
_T = TypeVar("_T", bound=Callable)
def inherit_signature(_function: _T) -> Callable[..., _T]:
return lambda f: f
# Requests for example has some long signatures (via typeshed).
class CustomSession(requests.Session):
@inherit_signature(requests.Session.post)
def post(self, url: str, *args: Any, **kwargs: Any) -> requests.Response:
...
And now for CustomSession.post the editor sees: def post(
self: Session,
url: str | bytes,
data: _Data | None = None,
json: Any | None = None,
*,
params: _Params | None = ...,
headers: _HeadersUpdateMapping | None = ...,
cookies: RequestsCookieJar | _TextMapping | None = ...,
files: _Files | None = ...,
auth: _Auth | None = ...,
timeout: _Timeout | None = ...,
allow_redirects: bool = ...,
proxies: _TextMapping | None = ...,
hooks: _HooksInput | None = ...,
stream: bool | None = ...,
verify: _Verify | None = ...,
cert: _Cert | None = ...
) -> ResponseDon’t forget to use ‘ConfigDict(frozen=True)’ absolutely everywhere!
> whether or not models are faux-immutable, i.e. whether __setattr__ is allowed (default: True)
Seems everything is “faux” in Python world when it comes to typing ;)
That said, there's also msgspec [1], which I've not used yet but plan to for my next project. It is supposedly quite a bit faster than Pydantic in de/serialization.