Python Standard Library changes in recent years
antonz.org
antonz.org
I know someone will blurt out “Haskell!” or whatnot, but I wish memoization was more first class in popular languages. It’s a great way to maintain performance without having to reason about the implementation details of your code. A decorator in the stdlib gets pretty close.
can you elaborate the problem a bit more please?
Simplest example: take any function you memoized with @functools.memoize. The more it's called, the the bigger the cache can grow. Run it on a stream of inputs and you're practically guaranteed to run out of RAM.
This obviously has nothing to do with referential transparency.
say you memoize a function that is referentially transparent, needs 10secs to finish its execution and return its result (eg. factorial). then the memoization would just cache the arguments plus the return value to reduce the 10secs into a simple comparison. the worst that could happen is a missed hit which would invoke the expensive function again - just like what would happen without the optimization. that isn't a big concern, as the function is referentially transparent (without side-effects). memoization lives in this perfect mathematical world where memory for maxsize entries will always be available.
but if your expensive function is referentially opaque, but still pure, then you can't only depend on the arguments of the function, but you need to take more context into account and the additional context requires a form of cache invalidation. the context can be counter-based (only re-evaluate every nth call), time-based (re-evaluate at most once every t time, throttling or debouncing), or both (only allow for at most n/t re-evaluations over time, rate-limiting) etc. in my book that isn't memoization anymore, it's just caching. and you need to define a policy that defines its invalidation.
if your stream of inputs exceeds the available memory, it lives outside the perfect mathematical world and i'd guess you have other problems than a speed optimization through memoization. you'd then need to consider space-efficiency and i would argue that memoization is the wrong kind of optimization for that problem.
var fm = add_memoisation_to(f); // used generics for c#/scala
Done this in javascript, scala, C#. Can't remember how I handled params, it's not a big deal.@dataflow: you build in cache limits, either by time or space, and instrumenting so you can get feedback (eg. hit rate, cache size, whether the cache turnover indicates thrashing). Very little code, a page or two.
[1] and if it doesn't, just use objects. Slightly less streamlined code but otherwise identical semantics.
Notice this is the same reason why you don't see a general-purpose "tree" or "graph" class in many languages' standard libraries [1]: you certainly can make one, but design has strong dependence on the actual problem you're solving, and it's hard to come up with a general-purpose one that's handy and useful. Of course, some people try anyway, and it occasionally gets used, but nowhere remotely as often as people need trees or graphs. Same thing here.
If you really want specialised caching policy then you can have that as an extra object/function passed in to the cache (edit: a predicate or predicate closure), but I've never needed it. LRU all the way for me and most others.
> Just look at common Python libraries (even the standard library itself) and count how many times they use @memoize anywhere.
Obviously! They are general purpose so they can't make any guess, and if they need caching that's up to the user, who can then add @functools.memoize in one line. That's what it's there for!
> And that fact is not widespread is not due to lack of knowledge of the technique.
IME it very much is.
> Notice this is the same reason why you don't see a general-purpose "tree" or "graph" class in many languages' standard libraries
But then caching is far simpler than graphs, and anyway from your link: "(Out of more than a dozen projects in various domains I have been involved in so far, I recount two which actually needed graphs.) So I guess there was no really big pressure from the Java community in general to have a Graph in the Collection Framework."
[0] https://docs.python.org/3/library/functools.html#functools.c...
Or - in PEP 628 - it's just a bit of a larf https://peps.python.org/pep-0628/
Perhaps this guy should grow a sense of humour.
Agreed!
I think while the experienced here might scoff at it, it still remains one of the most accessible and -fun- languages. Also I just adore iterators and any language without it is a no-go.
> I think while the experienced here might scoff at it, it still remains one of the most accessible and -fun- languages.
This is a shit attitude, and this comes from someone who mostly does C++ and Java at work, you know, "real languages" with braces and builds and whatnot. Python being easy and accessible isn't a bug, it's a goddamn feature.
I'm not sure dynamic typing is the fault by itself, I guess it's from a lack of compilation phase, as other languages such as Elixir seems to protect you from typing mistakes much more than Python does. My running theory is that dynamic typing is a sprectrum, and Python is on the side of "it's very easy to get it wrong" compared to other languages. I haven't used the recent versions with type hints, so it might've gotten better.
Also, I'm not a fan of how it feels. It is afraid of functional constructs: come on, itertools.accumulate is a simple reduce summing the elements. Why is Python so against map, reduce and ergonomic lambdas?
Like Go, it seems Python wants to be simple enough to be understood in large teams of varying developer skills, but unlike Go, Python is much harder to understand and debug in hairier codebases when you add metaclasses, weird backward compatible behaviours, monkey patching, typing woes and an often substandard stdlib.
I don't think this is true at all. There have been people against dynamic typing since it was first introduced, and their justifications haven't changed due to Python. There's nothing about Python's approach to dynamic types that is particularly egregious, in my opinion. In fact, since type annotations were added to the syntax in 3.6 (I think), there is now support for external tools to do some static type-checking (though these are not enforced by the language proper). If I were to point to an in-use dynamically typed language with dynamic type problems, I'd be pointing at JavaScript. The extend of implicit conversions is truly heinous (again, just my opinion).
> I'm not sure dynamic typing is the fault by itself, I guess it's from a lack of compilation phase
I'm not sure I understand this. For a language to be "dynamically typed", it must necessarily not check types during compilation no matter what, otherwise it would be "statically typed". So I don't see how adding a distinct compilation phase to Python's evaluation model would affect anything.
Also, Python does have a compilation phase; it just isn't distinct from the interpretation phase. There are no longer interpreters separate from compilers, as far as mainstream languages are concerned.
> Why is Python so against map, reduce and ergonomic lambdas?
The actual answer to this is that Guido van Rossum didn't want to make them first-class. That's the whole justification.
> Python is much harder to understand and debug in hairier codebases when you add metaclasses, weird backward compatible behaviours, monkey patching, typing woes and an often substandard stdlib.
Most of these sound like problems with dynamic types, honestly.
I'm pretty sure that honor falls upon PHP and perhaps Perl as well as JS.
> Python is much harder to understand and debug in hairier codebases when you add metaclasses, weird backward compatible behaviours, monkey patching, typing woes and an often substandard stdlib.
Indeed, Python (it's not just about dynamic typing, but more about how the underlying constructs of the language are exposed as first class citizen) makes it possible to write very bad programs, that no one would want to work with.
But this is also what allowed to build great libraries, that could abstract a lot from the user, particularly because of how much you can twist it to your needs. It's also great when writing tests.
My take is that there was a time were monkeypatching and doing this nasty stuff everywhere was common (early days of python becoming mainstream). Some libraries might still do this, but they are well tested and supported, so not an issue for them. But most python code being written now does not make much use of these features.
About functional constructs, I mostly agree with you, but the various comprehensions are a joy to work with, and provide enough for the most common usecases; the recent `match` construct also makes it possible to simulate ADT, which is great.
Finally, typing is very mainstream now, great libraries such as dataclasses or pydantic make classes with typed fields very easy to use, with minimal boilerplate.
You never saw a line of assembler, didn't you? You also just discounted 40 years of computer science (EDIT: I mean programming, CS has longer history of course) that happened before advent of Python. Good job.
Well, separating them is just as arbitrary. From the perspective of type theory, both kinds are just unityped. I was responding to claims about dynamic typing - and to that, assembler is relevant, as it uses the same typing discipline.
Moreover, as others noted, the division between static and dynamic typing fans probably started with Lisp sometime in the '50s. The arguments for and against each typing discipline are mostly the same as they were way, way before Python existed. Some technological breakthroughs did change some aspects of a discussion, but the debate is mainly philosophical and treated as a social activity more than an engineering one. That helped it remain frozen in time for so long.
In any case, Python is most assuredly not what gave dynamic typing a bad name. Instead, what gave dynamic typing a bad name were developers with a strong preference for static typing. And on the other hand, what gave static typing a bad reputation - and it also does have one like that - were developers who liked dynamic typing more.
That statement is ahistorical and is impossible to defend without intensive goalpost-moving.
Also is ZoneInfo basically stdlib version of pytz? Why not just put it under datetime.timezone like utc?
A good thing about Python is that pretty much every decision and rationale is documented by PEP, and this is no exception: https://peps.python.org/pep-0615/#using-datetime-zoneinfo-in...
In short the heavy implementation of zoneinfo disfavors a nested module, since it is not a norm in Python that importing A doesn't automatically make A.B available as well.
They optimised for, what, 20 people instead of optimising for thousands. I am not impressed at all.
“Optimising” here is quite a euphemism too. The omission of timezones from the standard library, alongside allowing timezone naive datetimes, has been a major implementation flaw for a long time.
One of such problems is to keep the database up to date, which is nothing but trivial. I'm very confident that the OS-level time zone data is the only database you can be reasonably sure that it's likely up to date, and different OSes expose (sometimes accidentally) different bits of data. Any other solutions are not very automated and tend to drift, I've seen systems where pytz or Java's TZUpdater neglected for a while. I hoped TZDIST [1] to be a thing, but it is not widespread enough, to say the least.
In this sense PEP 615 is not exactly the solution but a stopgap; it essentially admits that we can't exactly solve this problem, but we can at least include the existing half-solution into the batteries.
[1] https://docs.python.org/3/library/datetime.html#datetime.tim...
Also, you don't have to call is_active(), the object itself is convertible to bool. Not sure if that is an improvement.
graphlib is definitely underwhelming at the moment. A standard Python vocabulary for graphs would be nice.
- array: everyone uses numpy for the same purposes
- bisect: very niche feature, not a good fit for standard library
- glob: should probably be part of os
- graphlib: why is it part of the standard library? even if you need to work with graphs, chances are you have different requirements or data format
- shlex: again, not really necessary in standard library (unless it's used in shutil, I don't know)
- statistics: should be handled by scipy
And that's just the modules mentioned in this article. There are tons of other modules that are outdated, unused, don't fit into standard library: getopt, curses, urllib (everyone uses requests), xml*, html, tkinter.
What Python should be doing is deprecating those modules and/or moving them outside of the standard library. Not adding more of them.
For that, some of the included libraries really ease the pain of having to reinvent things like text-wrapping, special cases when iterating over things, creating TUIs, etc.
The more you read the docs, the more you appreciate the thought and effort of the Python developers in making things useful and convenient for everyone.
To me, the breadth of the standard library is one of Python’s main strengths.
This is what will be (eventually) removed: https://docs.python.org/3/library/superseded.html
> urllib (everyone uses requests)
iirc requests is actually a wrapper for urllib. And there are certainly many scripts using urllib directly to avoid having an out-of-stdlib dependency. The API is not that bad.
For example some modules are not what they seem to be; graphlib doesn't purport to be a general graph library and statistics doesn't mean to replace scipy. They are just groups for otherwise independent bits of code, named so that similar future code can be grouped together. In some cases you have a waaaaaay too high threshold for "unused" modules, for example [1] kept getopt because it mirrors a popular C library.
It's an excellent reason to keep it as a separately maintained module. Having in the standard library two modules for the same purpose just adds to the confusion.
Having them included by default is much more convenient; I may want to do some very simple stats without needing all of scipy, etc.
My usage is generally lists for small stuff where I don't care about performance, and pandas/numpy for analysis. So maybe array is either if I don't want numpy, or if I am doing something real and large and list performance would truly be a problem?
Also, I looked for a PEP to answer these questions and couldn't find it. (The one I found was very old). Or any stats on performance improvements?
I also think having basic statistic computations is useful. There are people who may want to calculate simple statistics without adding a whole dependency.
I don't want a dependency on NumPy, especially as I don't need any other feature of NumPy beyond storing a homogeneous, resizeable vector of simple data types.
No, thanks.
It still is used in situations where numpy doesn't fit (serialization of data, etc)
Agree with shlex and glob
> statistics: should be handled by scipy
It was just added. I guess for people who want to import numpy to calculate the mean of an array? (please don't do this - you don't need to do this)
getopt is a needed evil and I think requests uses a lot of builtin stuff
Python does deprecate some modules once in a while
For one thing, you can pry Tkinter from my cold dead hands. For another, I used array just last week.
This whole "no one is using it so throw it away" culture is part of the improving-to-death of the modern Python ecosystem. I think it must come from the Python 3 fiasco. It's become fashionable in a subset of Python community to aggressively deprecate working code.
Being able to provide solid globbing options to users in any scenario that is 'path-y' can provide massively more power to them, in lots of different cases I'd say.
The name isn't great, though. It actually offers double dispatch, not single dispatch! `divider.divide(10, 2)` first dispatches on `divider`, then on `10`.
As so often with Python, I'm left wondering why they didn't follow through and properly complete the design, in this case perhaps offering multiple dispatch (multimethods) instead of just double dispatch. Casting an eye over the functools code, it looks like they have all the necessary C3 infrastructure in there.
That code looks ugly as fuck though
key = operator.itemgetter("name")
people = sorted([p1, p2, p3], key=key)
What's the advantage of using operator module instead of just key=lambda x: x['name']
? import operator
n1 = operator.itemgetter('name')
n2 = lambda x: x['name']
x = {'name': 'abcdef'}
import timeit
print(timeit.timeit('n1(x)', globals=globals()))
print(timeit.timeit('n2(x)', globals=globals()))
I'm so surprised because I went and read (what I thought was) the source for operator.itemgetter, which boils down to a wrapper object around the plain version.It turns out that this is not the source code. As with so many other things in Python, there's a special case you have to know about: there's a separate C implementation of the operator module. :-/
I actually kinda like that Python often maintains 100% compatible pure-Python implementations of various C modules. It's definitely niche, but I have sometimes found myself trying to work out what a module is doing (e.g. with the fancy new PEG parser) and I find it far easier to read Python than C.
PyPy can do it with its jit because it supports invalidating compiled code when such changes are made.
... except for the stack trace:
>>> f = lambda x: x["A"]
>>> f({"A": 3})
3
>>> f({"B": 3})
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "<stdin>", line 1, in <lambda>
KeyError: 'A'
>>> import operator
>>> f = operator.itemgetter("A")
>>> f({"A": 3})
3
>>> f({"B": 3})
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
KeyError: 'A'
So if 'x' implemented its own getitem and it inspected the stack (eg, how Pandas uses the 'level' parameter in 'eval' and 'query' methods to get locals() and globals()), then this change would cause different results.I think you are right.
operator.countOf() is perhaps the hardest to reimplement in lambda, as it's more reliable than lambda x: x.count() - for example it will work for dictionary.values() as well.
people = {
"Diane": 70,
"Bob": 78,
"Emma": 84
}
keys = people.keys()
# dict_keys(['Diane', 'Bob', 'Emma'])
keys.mapping["Bob"]
# 78
EDIT: I actually misunderstood this; see comments. The keys.mapping field does not point to the original dictionary. I guess this is actually just about consistency with the way other mappings work? Not so useful (to me) after all.Or is this about read-only access? There are better, more explicit ways to accomplish that.
people = {
"Diane": 70,
"Bob": 78,
"Emma": 84
}
people.get("Bob")
# 78
The common usecase is: for key in people.keys():
people.get(key)
vs keys = people.keys()
# dict_keys(['Diane', 'Bob', 'Emma'])
for key in keys:
keys.mapping[key]
What is the advantage?EDIT: apparently not:
>>> keys.mapping
mappingproxy({'Diane': 70, 'Bob': 78, 'Emma': 84})
In this case I actually don't know what the use case is :shrug:https://github.com/python/cpython/issues/85067#issuecomment-...
It seems a bit vague: « Exposing this attribute would help with introspection, making it possible to write efficient functions that operate on dict views. »
e.g.
'abcdef'.split(2) == ['ab', 'cd', 'ef'] # not the same as .split("2")
[1, 2, 3, 4, 5, 6].split(3) == [[1,2,3], [4,5,6]]
Or one can use the `grouper` from the recipe
https://docs.python.org/3/library/itertools.html#itertools-r...
Pairwise would give you ['ab', 'bc', 'cd', 'de', 'ef']. It's not a split.
[1]: https://docs.python.org/3/library/stdtypes.html#int.bit_coun...
Edited, more info: AMD's Barcelona architecture introduced the advanced bit manipulation (ABM) ISA introducing the POPCNT instruction as part of the SSE4a extensions in 2007. Intel Core processors introduced a POPCNT instruction with the SSE4.2 instruction set extension, first available in a Nehalem-based Core i7 processor, released in November 2008.
https://www.researchgate.net/figure/Census-transform-and-Ham...
- A bitset index
- Bitflags (e.g. to represent logging levels)
- Feature flags
`bit_count` (aka `popcount`) gives you an efficient way to figure out how many bits within those structures are set and not set.
I guess it's included as an easy way to access the fast hardware supported implementation, on architectures where it is available.
This article has a lot of interesting examples. https://vaibhavsagar.com/blog/2019/09/08/popcount/
(0.2).as_integer_ratio()
# (3602879701896397, 18014398509481984)
# oopsie