Learn by reading code: Python standard library design decisions explained
death.andgravity.com
death.andgravity.com
I'm working on a slightly similar project motivated by this problem: how do you learn from established open-source projects - most interesting ones are too big, too complex, to hard to get started with.
So I'm compiling a collection of interesting code examples from open-source projects and explaining/annotating them. I'm trying to pick examples that are:
- Taken from popular, established open-source projects.
- Somewhat self-contained. They can be understood with little knowledge of the surrounding context.
- Small-ish. The can be understood in one sitting.
- Non-trivial.
- Instructive. They solve somewhat general problems, similar to what some other coders on some other projects could be facing.
- Good code, at least in my opinion.
I'm planning to share it in a few weeks.
It's empty for now, but I included a few examples [1] of what I'm going to annotate and add there.
Specifically how they are untarring each container layer and creating a chroot jail to run commands inside is fairly self-contained and interesting.
My answer for this has been to scroll all the way through the git history of a repo structurally similar to something I want to do and read each of the first month-or-so of commits.
I'd agree that pathlib is a good solution, and that it's an object-oriented solution, but I'm hesitant to call it a good object-oriented solution. It uses some object-oriented hackery to work as elegantly as it does. I approve of this decision, but it leads to some weirdness such as the inability to subclass `pathlib.Path`. Instead you need to subclass `type(pathlib.Path())`, as the method resolution order changes once it is instantiated:
>>> from pathlib import Path
>>> Path.__mro__
(<class 'pathlib.Path'>, <class 'pathlib.PurePath'>, <class 'object'>)
>>> type(Path()).__mro__
(<class 'pathlib.PosixPath'>, <class 'pathlib.Path'>, <class 'pathlib.PurePosixPath'>, <class 'pathlib.PurePath'>, <class 'object'>)
In any case, weirdness like this can make for an even better learning opportunity! >>> with open("pickle.pkl", "rb") as f:
... pickle.load(f)
Traceback (most recent call last):
File "<stdin>", line 2, in <module>
File "/usr/lib/python3.9/pathlib.py", line 1074, in __new__
raise NotImplementedError("cannot instantiate %r on your system"
NotImplementedError: cannot instantiate 'WindowsPath' on your systemTake Rust's std::iter::repeat_with
https://doc.rust-lang.org/std/iter/fn.repeat_with.html
There's a declaration of this standard library function to make an iterator from a closure, then there's a working code sample which uses it and a further example showing a closer to real-world usage.
But top-right is a [src] link, you can go in one click from the documentation of the function call, to the actual implementation which is used (except where Rust just says this is a compiler intrinsic and then you need to burrow into the details of the actual compiler).
Big if true!
N.B. I've been programming in Python since it was version 1.3. Forgive me.
Why doesn't someone make a plain language annotation -- for beginners -- to translate the official documentation?
I ought to do more of it but I think this ought to become a thing - a sort of nutshell guide to different modules.
I think this should be something between the manual and pymotw
if iter(data) is data:
data = list(data)
n = len(data)
if n < 1:
raise StatisticsError('mean requires at least one data point')
T, total, count = _sum(data)
assert count == n
return _convert(total / n, T)
The _sum function is a little more involved, but not appreciably so.In general though, a literal translation of a formula is not always great in terms of numerical stability/error. For example, you might have learned to calculate the variance as mean(X^2) - mean(X)^2, but this can lead to a huge loss of precision that more complicated approaches avoid.
> if iter(data) is data:
Wouldn't it be cheaper like `if type(data) is iter:`?
And why convert `data` to a list at all, to check for length 0, given that `_sum(data)` will return the count?