The return of lazy imports for Python
lwn.net
lwn.net
Manual lazy import meaning:
def uses_foo(x):
import foo
return foo.bar(x)
Manual lazy import sucks because:- it's just ugly, I like imports all at the top so I can see all the deps
- bad for static analysis
- performance hit every time the function is called
Eschewing lazy imports has several problems:
- you always pay the execution cost, even if you don't use it
- also bad for static analysis and testing, since you have to eat the import time even if the code block you want to test doesn't execute the expensive path
- sometimes you need lazy import to avoid circular import errors
It's too bad that the main impediment to this is existing code which relies on side effects. Import-time side effects are an absolute pain in the ass. Avoid it at all costs.
It’s a more advanced type system than Go and Java
It’s very trivial to have something pass the type hinting checks, but be completely different at time of use. Even if you hold on to a single instance of an object, it’s easy for any thing to monkey patch it.
Python modules are always singletons, regardless of where they are imported. "import foo" inside a function will only import the module once (=global effect) but bind the name every time (=local, but cheap, effect).
Besides, it’s Python. It’s not going to be super fast anyway. That extra check is never going to show up on a perf trace.
$ cat spam.py
import sys
print(f'__name__ = {__name__!r}')
for (name, mod) in sys.modules.items():
try:
if 'spam' not in (mod.__file__ or ''):
continue
except AttributeError:
continue
print(f'sys.modules[{name!r}] = {mod!r} @ 0x{id(mod):x}')
import spam
$ python3 -m spam
__name__ = '__main__'
sys.modules['__main__'] = <module 'spam' from '/home/jwilk/spam.py'> @ 0xf7d86ed8
__name__ = 'spam'
sys.modules['__main__'] = <module 'spam' from '/home/jwilk/spam.py'> @ 0xf7d86ed8
sys.modules['spam'] = <module 'spam' from '/home/jwilk/spam.py'> @ 0xf7caef78
More generally, Python is happy to import the same file multiple times as long as the module name is different.For example, if there's eggs/bacon/spam.py and you add both "eggs" and "eggs/bacon" to sys.path, you will have two different modules imported after "import bacon.spam, spam".
def uses_foo(x):
global uses_foo
import foo
uses_foo = foo.bar
return foo.bar(x)
While that performance suckage improves, the other suckages get worse.Tried some poor-man’s debugging and never hit a breakpoint on the first significant line of code… took a while to figure out as it was my first Python project.
It almost feels like Python needs a scripting and non-scripting mode, or some kind of warning logging “you did everything wrong”
Personally I am betting on JavaScript being the first dynamic and JITted (or incrementally compiled) language that can attain the AOT compilation without massively bloated binaries.
> (or even Java/.NET style lowered bytecode binaries).
Traditional AOT like you see in C/C++/Rust/Go/Zig would be able to treeshake and eliminate redundant codepaths, the binaries are all fairly small with minimal startup overhead.
Another counterexample to AOT compilation being a feature of application languages is Java and C# -- both are clearly application languages, with strong focus and presence there, but both are interpreted. Although I can argue that it depends on the terminology regarding "compile", whether transpilation to .class bytecode can be called compilation (and everyone do call it "compilation", but if it really was, there would be no point in real AOT compilers for Java/C#).
Although I found one feature to very strongly correlate -- static typing. For example, Julia is AOT compiled, but still dynamically typed. And coincidentally, it was designed for scientists, who need scripting much more.
Another observation is how easy it is to write unit-tests (for canonically written code). Overloading "import" statement doesn't help there at all.
Our fix for this was to relegate stateful imports to files with "bootstrap" in the name, which the lazy loader stub would allow to be eagerly imported. Moreover, any imports listed in such "bootstrap" files would then be eagerly imported. But that's it (at least wrt the code belonging to our codebase); no one else, not even the children of the bootstrapping modules, were exempted from lazy loading.
This allowed for a centralization of all the stateful import effects. If you tried to write stateful code outside of bootstrapping, it simply wouldn't run. (You could hypothetically hack around this. But it would be rather obvious, and I'm perfectly ready to revert pull requests that try.)
Maybe you folks could try some variation of this in Python.
# jobs.py
@scheduler.register_job(when=“daily”)
def batch_job(): …
This adds the job to the global scheduler. # models.py
class Model(DBModel):
def on_change():
# do stuff
This adds the model to the ORM’s list of modules via a metaclasss and registers its hooks. # plugin.py
class SomePlugin(PluginBase): …
Same thing. The code to load plugins like this is just importing the module and letting the metaclass do the work.It is absurdly easy in Python to end up with a circular import situtation, where no real circular dependency exists. E.g. you can't have A.a1 -> B.b1 and B.b2 -> A.a2. So, you are forced to layout your code in some quite awkward ways.
if TYPE_CHECKING: import WhateverClass
https://docs.python.org/3/library/typing.html#typing.TYPE_CH...
(And some types are defined in the typeshed so only exist to be imported during type checking; eg the type checker lib itself is a dependency in this case)
if False:
import blah
unironically as good design and more than a necessary evil until a long-term solution emerges then we’ve jumped the shark. from typing import List
class Alpha:
@staticmethod
def doit(b: "Beta") -> List["Beta"]:
return [b]
class Beta:
@staticmethod
def doit(a: "Alpha") -> List["Alpha"]:
return [a] import gamma
def doit(bar: "gamma.Gamma"): ...
In Java my answer to circular deps is the introduction of an interface that the concrete types can implement but then breaks the cycleYou don't even need functional import syntax, but as TFA notes this comes at a cost as it has to invoke the entire import machinery, which it can only skip to an extent (once a module is loaded and cached) as import hooks can have odd behaviours.
For B), easy enough to run one of many linters to detect this case and have people write less bad code.
A) is way more subjective and can be fixed in many ways.
With the many more Python coders these days with less coding experience, personal feeling is please stop throwing these production issue causing features in that I have to fix. Glad the PEP is rejected.
Old programmer wisdom is to load all your configs and assumptions as early as possible to eliminate a whole space of problems with your code, making faster and easier to read/reason about later.
I’ve seen a moderately sized (~300k LoC) Python CLI project that had a horrendous, anger-inducing startup time until they switched to the lazy import approach basically described/standardized by PEP 690 and the improvement was massive.
It doesn't have to be like that though -- look at geohotz's tinygrad library for exampe: well tested, well written, and can do most of the things the bigger ecosystem libraries do.
Conversely, agreed, it's not THAT hard to convert an individual project, just being unable to handle third-party, so import-shaming some top projects can likely go far. Features that enable framework providers to empower others are great... And none is needed here. It's messy to push through type checking, but doable.
Feel like we are reliving when js got a more static module system. A lot of kicking and screaming, and still issues, but a lot of good came out of that.
https://docs.python.org/3/library/importlib.html#implementin...
The most egregious is when you just want to display the help/usage text and not actually execute any code at all. Instead, you have to either manually lazy import (what I usually do now), or eat huge startup costs each time you screw up the command syntax.
Replaced it with a 20 line c extension that imported basically instantly.
Xonsh shell is amazing.
>>> import sys
>>> len(sys.modules)
83
>>> import numpy as np
>>> len(sys.modules)
220
>>> sum(1 for k in sys.modules if "numpy" in k)
94
so people can write a one-liner like: >>> np.polynomial.chebyshev.Chebyshev([0,1,3])(np.linspace(-1.0, 1.0, 5))
array([ 2., -2., -3., -1., 4.])
without having to import np.polynomial.chebyshev.Chebyshev first.This API design requires importing most of NumPy at startup, which has a cost they didn't consider so important because their users are primarily doing long-term computing and notebook-style development, where startup cost is relatively small.
I've complained about this because I live in the short-lived program world, where it's annoying to have a 0.1 second import overhead if I only need one function from NumPy:
py310% time python -c 'pass'
0.025u 0.006s 0:00.03 66.6% 0+0k 0+0io 0pf+0w
py310% time python -c 'import numpy'
0.142u 0.292s 0:00.14 307.1% 0+0k 0+0io 0pf+0w
As I understand it, SciPy wants a similar API design goal, but has a lot more packages. They've developed lazy imports to try to have the best of both worlds.> For command line utilities, it seems like you're going to need to load the module at some point or another regardless (if you're actually using it)
Thing is, you might not actually use it. If the command-line tool uses subcommands, each different subcommand might need only a subset of the full set of packages.
Perhaps only one of the subcommands uses NumPy, while for 95% of uses, NuPy isn't used at all.
As the discussion for this feature points out, this can be addressed by only importing when needed. (One of the reasons I've started using click over argparse is click does more of this separation for me.) However, it's somewhat fragile, in that it's easy to add an rarely-needed expensive import at top-level without noticing it, and requires some non-standard tooling to detect issues, like the non-predictability you mentioned.
I personally want something like the lazy-/auto- importer in my package, so I can reduce the two step process. My last package released used module-level getattr functions, which gets me mostly there, except for notebook auto-completion of the lazy wrappers. (It works in the command-line shell though.)
I can't import everything on startup because parts of my package depend on third-party packages, which might not be installed. I instead want to raise an ImportError when those lazy objects are accessed. Plus, one of the third-party packages is through a Python/Java bridge, which has its own startup costs that I want to avoid.
I don’t think lack of facilitation skill is an issue; its a deliberate policy choice.
There's variations of degree, but probably not. Part of being production-ready is stability.
The are a lot of reasons not to introduce it, but 'what ifs' at a company like ours could be devastating. I still think proper precautions can be taken, but it is harder for me to say that I would just say yes if I was in his shoes.
You need developers who care about fast, clean code to fix the issue. Those kind of developers usually don't fare well in the Python swamp, so it won't happen.
argparse with subcommands generally requires specifying all of the options for all of the subcommands, even if you only want one subcommand.
These in turn may require importing subcommand-specific modules, to handle things like the right 'type' handler in an an add_argument() parameter. This callback function might, depending on the input value, select one from a dozen different additional packages.
It's possible to avoid this, by deferring argument->type processing until later, and having a single large module containing all of the help strings and epilogs, though this will separate your argparse code from your subcommand code, and in general make things more complicated. I did this for a while.
Alternatively, you can create your own subcommand dispatch system using an nargs="?" to get the subcommand and an nargs=argparse.REMAINDER to capture the rest of the flags, to pass to a new ArgumentParser, and develop a top-level --help replacements. I tried this too.
I've since decided to use click, which does a better job at compartmentalizing at least this level of subcommand imports.
mylib = None
in the global scope and then
global mylib
if mylib is None:
import mylib
in your function to avoid the extra function call.E.g first the local dict, then the global and then the modules dict, instead of just the global+modules
import mod
is the equivalent of
mod = sys.modules.get(name)
if not mod:
mod = sys.modules[name] = load_package(name)
It’s really low cost, as long as you’re not doing it in a hot loop, it’ll be very low to no impact.