How the Python object system works
tenthousandmeters.com
tenthousandmeters.com
- What Python objects and types are and how they are implemented.
- What slots are and how they determine the behavior of objects.
- How slots are related to special methods.
I'd be glad to get your feedback and answer your questions. Thanks!I don't have any questions. I just wanted to let you know I think you do a fantastic job writing these up. The thinking process is very clear, and the resource links are well placed for followup research. High quality stuff, for sure.
Thanks for making these. I look forward to future installations.
In the spring I’ll be teaching a graduate-level university course on how the Python compiler and tooling works, so I’ll be showing them your articles.
Both of your examples bring very little meaning from their English definition. You just know a second definition in the context of programming, which is why you can deduce what it does.
The same applies to the people that don't fluidly speak the language.
So I agree, you don't have to speak fluent English to be able to read and write programs. Knowledge of English grammar won't help (except for reading comments – if they exist...), and the vocabulary is small, specialized, and has little to do with colloquial language.
[1] Or should it be "...even stronger"? I don't think so, because it refers to the verb "goes", so it should be an adverb, right?
Here in Europe is quite common to be faced with programs where everything written in-house uses the local language, including comments, only the third party libraries use English.
I already had a couple of consulting gigs where what got me the deal was my knowledge of the specific (human) languages being used.
I've encountered programmers who are not fluent in English (e.g., some of the audience of non-English StackOverflow) but they don't seem to be a curious type (for whatever reason)
(Before I get bashed for this, I speak 3 languages well and 4 more poorly, English is my third. It really isn’t a difficult language to learn when all American media and most of computing uses it.)
It is ok that my comment above is downvoted but It would nice to hear any counter arguments (it is hard for me to imagine a passionate programmer who can't learn English).
Language learning is difficult and your success with one language has a lot to do with what you speak natively. I've heard that Italian speakers for instance can pick up English with relative ease compared with someone who grew up speaking Japanese.
People also have different priorities in life. Past high school and university you've got a career and possibly a family to attend to. Finding enough time to learn another language would be a big ask for many.
The link between language learning and curiosity sounds tenuous at best. I would also add that having access to media in a particular language doesn't mean a lot unless you're prepared to put some effort in. I know expats who've lived in Japan for 20 years and can only speak a few words of Japanese despite being surrounded by it everyday.
They said that there might be because they know people who are interested in the subject but don’t speak English. So I don’t know why you are asking the question.
Not sure you're at all interested in going in this direction, but I (and likely many people) would be really interested in a similar explanation of NumPy internals.
As for python internals, https://snarky.ca/ has some great short high level “how does this construct work” text explainations.
class C:
def __init__(self):
self.a = 1
def b(self):
print(f'b({self})')
@property
def c(self):
print(f'c({self})')
o = C()
The descriptor protocol is a generic way to enable `o.a` to access an attribute of `o`, `o.b` to access an attribute of `C` and bind `self` to `o`, and `o.c` to access an attribute of `C`, bind `self` to `o`, and call it.Do any other dynamic languages use a similar protocol for attribute access?
Ruby, for example, doesn't need to - it can use a simpler approach because `o.x` never refers to an attribute of `o`.
JavaScript's solution is less elegant - the way it sets the value of `this` is infamous, and AFAIK properties are treated as a special case rather than built on top of a generic protocol.
But the cpython folk make understandability an explicit goal of the codebase which tends to preclude things like fancy optimisations or JITs
myObj.MyMethod()
would be implemented as (this is from memory and I'm rusty):
myObj is fetched from the local scope array by array-index.
Dictionary lookup on myObj failing over to dictionary lookup on myObj.__class__ to find the method MyMethod(). Or was it one merged dict? Whichever. All __slots__ does is mean that myObj doesn't need its own dictionary, it still has to go to __class__ for a dictionary hit, which can even be overridden if they've changed __getattr__ or __getattribute__.
Originally cpython even used deliberately-high-collision hashtables for lookups for reasons I no longer remember (faster sort?).
MyMethod() instantiates a new Method object that stores the underlying class-function and the "self" parameter, but the object is pooled, so it's quick. Still, it's creating a reference that will have to be collected in a moment.
Then we actually invoke the damned function.
Then we decrement the refcount on the Method object, which drops it to zero and it is destroyed (returning it to object-pool).
Obviously, this is from my memory and it's been over a decade, so I might be getting details wrong, but it was surprised that there were so many unoptimized layers involved in resolving a method. That even with __slots__ dictionary hits were unavoidable, and that method invocation involved instantiating an object (from a pool, but still).
> Obviously, this is from my memory and it's been over a decade
I would not remember those details if I read them last week let alone a decade ago.
> myObj is fetched from the local scope array by array-index
That's true if myObj is a known local variable (assuming you're inside a function or class block--in global scope there is no "local scope array" to begin with). But if it's not, myObj has to be looked up in the dictionary of global variables, which is slower than the fast local array indexing. (And there is also the nonlocal keyword, which further complicates the lookup since enclosing non-global scopes also have to be included.)
> Or was it one merged dict?
No, it's separate. And it's further complicated by descriptors; first the lookup needs to check if MyMethod is a descriptor (such as a property) on the class, and if it is and the descriptor is a data descriptor (i.e., has a setter method), it overrides the lookup in the instance dictionary.
> All __slots__ does is mean that myObj doesn't need its own dictionary, it still has to go to __class__ for a dictionary hit
Yes, __slots__ is an optimization to reduce memory consumption, not to increase speed.
> surprised that there were so many unoptimized layers involved in resolving a method
They can't be optimized in the general case without sacrificing the dynamic attributes of the language, which would defeat the purpose.
Optimizers like PyPy focus on optimizing these layers in the special cases where particular dynamic attributes aren't being used in particular parts of the code. Cython, which does as much static analysis at compile time as possible to enable eliminating the extra lookups when they're not going to end up changing anything, is another example of the same idea.
No, it doesn't. There are two separate opcodes involved: LOAD_FAST loads a local variable from the local array based on its index; LOAD_GLOBAL loads a global variable based on the global dictionary lookup.
It can increase speed though in practice. Less memory means less management and better cache usage. I've nearly doubled the speed of stream reconstruction from packet capture with a lot of slots usage. (Huge amount of tiny objects)
In view of a follow-up comment about slots, this should be clarified: an class with __slots__ does not need to do a dictionary lookup for slot attributes; those are looked up by index into an array, just as local variables in a function are. The dictionary lookup is only done for attributes that aren't listed in __slots__.
Yes, this is a fair point. Still, accessing the slot does not require a dictionary lookup, as it would for an ordinary instance attribute, which was the main point I was trying to make.
> there're ways to optimize that lookup
The way CPython does this, if you can call it an "optimization", is to implement the lookup as a data descriptor, which directly accesses the slot array location by index. (The namedtuple implementation correspondingly implements accessing the attribute as a read-only, non-data descriptor that directly accesses the appropriate tuple location by index.) Quite possibly the fact that the descriptor lookup comes before anything else in the attribute access code is considered "optimization" enough for this case.
I'm sorry but mapping a string (slot name) to an index [in an overdynamic language like Python] does require a dictionary lookup. It's just this dictionary is located in the class, not in each instance.
> The way CPython does this, if you can call it an "optimization"
The usual way to optimize lookups-by-name in dynamic languages is using (inline) caches. AFAIK, CPython now does that too.
Yes, you're right, I wasn't clear enough. What I meant to say was that accessing the value of the slot attribute (to either get or set it) on the instance does not require a dictionary lookup, just an array access. But of course finding out that the string (attribute name) is the name of a slot and getting the slot index does require a dictionary lookup (on the class, as you say).
I'd be really curious about this in other languages as well, such as Ruby and JavaScript.
This part was always magical. It seems to free up the local memory usage of the program, but it doesn’t seem to release the memory back to the operating system, until some type of triggered event.
From this, I guess CPython opted for simplicity instead of implementing something like a memory-usage monitor.
[0] https://realpython.com/cpython-source-code-guide/#garbage-co...
That's up to the C runtime's memory allocator. Modern memory allocators don't typically request a new chunk of memory for every malloc() call -- instead, they allocate a single large region of memory at a time and carve that up as needed. This is massively more efficient (system calls are expensive), but also means that those regions can't be released to the OS until all allocations in them are gone.
The Python/C Reference Manual has a great section on memory management [3].
[1] https://docs.python.org/3/c-api/typeobj.html#c.PyTypeObject....
Lessons learned: Don't ever look at the python vm, if you want to look for a well designed VM. Even Ruby is miles better there.
Are there any good guides or maybe an introductory tutorial somewhere on how to get involved?