Multiversion Python Thoughts
lucumr.pocoo.org
lucumr.pocoo.org
If I have a special case and I need to do this today, I’m not blocked from doing so (I’ll vendor one of the dependencies and change the name) - certainly a pain to do, especially if I have to go change imports in a c module as part of the dep but achievable and not a blocker for a source available dependency.
However, if this becomes easily possible, well why shouldn’t I use it?
The net result is MORE complexity in python packaging. More overheads for infra tools to accomadate.
The alternative is just cheating: ignore requirements, install packages and hope for the best. Alas hope is not a process.
You'd obviously need to have tests for both targets, possibly using a flexible test runner like `nox` to setup separate test env for each target.
In C, C++, maybe Java, you would at least be able to link A and B with their own private copies of C to avoid conflicts reliably with standard mechanisms rather than unreliably with clever magical tools.
import version1
import version2
try:
version2.load_data(input)
except ValidationError:
version1.load_data(input)
I'm sure people will abuse it, but the idea doesn't seem terrible on the face of it to me. import version2
try:
version2.load_data(input)
except ValidationError:
import version1
version1.load_data(input)
(Assuming version in name like this example version2 or lib2 etc.)With a tool like rope, I feel fairly confident I can refactor a source available, pure python dependency, pretty quickly.
Where I get less comfortable with my idea is that not every dependency has source available (e.g. db2 database driver as an example).
Another case is where some deps which have source available but the python module dependency is in C/C++/Rust - e.g. scipy
Of course, you can change the import name between versions. That's one of the upsides of not tying the import name to the distribution name, and many real-world projects actually do this as part of a deprecation cycle (for example, `imageio` has been doing it with recent versions, offering "v2" and "v3" APIs). But in the general case, you'd have to change it with every version (since your transitive dependencies might want different minor/patch versions for some obscure reason - semver is only an ideal, after all), which in turn means your users would always have to pin their dependency to the exact version described by the code.
Instead, we jump through hoops with our hair on fire to manage complexity.
People.
I'm an advocate for this style of library/dependency development, unfortunately in my experience the average dependency doesn't have the discipline to pull it off.
Code changes, and it's kinda silly to expect interfaces to be locked in place as that'll stifle development for even small-ish features. Does that mean every minor version or commit will change fundamental or large parts of the codebase? Probably not, but it's a sliding scale and people seriously need to find something better to do than writing Yet Another Python Package Manager.
I use the term "we" loosely here ofc in the context of this mini-rant.
If it really can't be done then you aren't really shipping a library that is meant to be depended on.
Imagine trying to do things like import pyarrow14, then if that failed try pyarrow13, etc. Additionally, python doesn't have a good way of saying "I need one of the following different libraries as a dependency"
One:
Library A takes a callback function, and catches “request.HttpError” when invoking that callback.
The callback throws an exception from a differing version of “request”, which is missing an attribute that the exception handler requires.
What happens? How?
Two:
Library A has a function that returns a “request.Response” object.
Library B has a function that accepts a “request.Response” object, and performs “isinstance”/type equality on the object.
Library A and library B have differing and incompatible dependencies on “request”.
What version of the request object is sent to library B from library A, and how does “isinstance”/“type” interact with it?
Both of these resolve around class identities. In Python they are intrinsically linked to the defining module. Either you break this invariant and have two incompatible/different types have the same identity and introduce all kinds of bugs, or you don’t and also introduce all kinds of bugs - “yes, this is a request.Response object, but this method doesn’t exist on this request.Response object”, or “yes this is someone’s request.Response object, but it’s not your request.Response object”
Getting different module imports to succeed is more than possible, getting them to work together is another thing entirely.
One solution to this is the concept of visibility, which in Python is famously “not really a thing”. It’s safe to use incompatible versions of a library as long as the types are not visible - I.e no method returns a request.Response object, so the module is essentially a completely private implementation detail. This is how Rust handles this, I think.
However is obviously fucked by exceptions, so it seems pretty intractable.
I think if Python were to want to go down this path it should be isolated to explicit migration cases for specific libraries that want to opt themselves into multi-version resolution. I think it would enable the move of pretty core libraries in the ecosystem in backwards incompatible ways in a much smoother way than it is today.
Go and JavaScript have type systems and idioms far more amenable to this kind of thing (interfaces for Go, no real type identity + reliance on structural typing for JS) and rely a lot less on the kind of reflection common in Python (identity, class, etc).
I guess there are some use cases for this, I just feel that the lack of ability to enforce visibility combined with the “rock and a hard place” identity trade-off limits the practical usefulness.
It seems a lot more impactful with Python due to type equality being core to how exceptions are handled, even if there are similarities.
Say a function returns `fmt.Errorf("Error while doing intermediate operation: %w", lowerLevelErr)`, where `lowerLevelErr` is `ModuleBError`. Then, if you do `if _, ok := err.(ModuleBError) {...}`, this will return false; but if you do `if errors.Is(err, ModuleBError)`, you will get the expected true.
Regardless, the core problem would be the same: if your code can handle moduleB v1.5 errors but it's receiving moduleB v.17 errors, then it may not be able to handle them. This same thing happens with error values, Exceptions, and in fact any other case of two different implementations returned under the same interface.
You even have this problem with C-style integer error codes: say in version 1.5, whenever you try to open a path that is not recognized, you return the int 404. But in 1.7, you return 404 for a missing file, but 407 if it's a missing dir. Any code that is checking for err > 0 will keep working exactly as well, but code which was checking for code 404 to fix missing dir paths is now broken, even though the types are all exactly the same.
Sure, but that just means your dependency was not really internal. Errors are API too.
It follows the approach of "objects from one version of the library are not compatible with objects of the library" mentioned above, and results in a compile time error (a potentially confusing type error, although the error message might call out that there's multiple library versions involved).
It should be safe to use multiple versions of the same library, as long as they are used as private dependencies of unrelated dependencies. It would require some tooling support to do it safely:
1. Being able to declare dependencies are "private" or "public".
2. Tooling to check that you don't use private dependencies in your interfaces. This requires type annotations to gain some confidence, but even then, exceptions are a problem that is hard to check for (in Python that is).
In compiled languages there are additional compilications, like exported symbols. It is solveable in some controlled circumstances, but it's best to just not have this problem.
Herein lies the issue: in this context exceptions can be thought of as the same as returns. So you actually need to catch/handle all possible exceptions in order to not leak private types.
Also what does “except requests.HttpError” do in an outer context? It checks the class of an exception - so either it doesn’t catch some other modules version of requests.HttpError (confusion, invariants broken) or it does (confusion, invariants broken).
The requests HTTP exception contains the request and response object. Wrapping that would be a huge pain and a lot of code.
1. pip install libfoo==1.x.x
2. pip install libfoo==2.x.x --target ~/libs/libfoo_v2 # vendor libfoo v2
3.
import sys
import libfoo
original_sys_path = sys.path.copy()
sys.path.insert(0, '~/libs/libfoo_v2')
import libfoo as libfoo_v2
sys.path = original_sys_path
There are caveats of course. But works for simple cases.
I am using python as my main language these days, coming from JS and C++ and a bit of rust. The biggest problem I face is that the tools (editors, mainly) don't support the basic packaging tools.
I use venv for everything, but when I try to use something else, I almost always get bitten.
For example, asdf. I tried to use this tool. So awesome! It works great from the command line. But, when I try to use zed, it cannot figure out what to do, and I cannot find references in the zed github repository on the right way to setup pyproject.toml.
And, emacs. Will uv work within emacs? Each of these packaging tools (and I'm thinking about the long history of nvm (node version manager), brew, and everything else) makes different assumptions about the right way to modify the path variable, or create aliases, or use shims (the definition of which varies with each tool) and I'm sure I'm missing other details.
Does uv do the right thing mostly? I will say my experiences with python and the tooling has been more frustrating than the tools for JS. I use pnpm and it just works, and I understand the benefits. And, I can survive with just npm and yarn. But, to me, it is saying a lot that the python tooling feels more broken than JS. I mean, I lived through the webpack years, and I'm still using JS and have a generally favorable opinion of it.
I hold the same opinion as you. Python packaging is awful. But uv managed to just make it work.
I removed asdf because I could not get it to recognize the pyproject.toml file. This is working in harmony for you?
With Zed and a venv (from python -m venv .venv for example), Zed properly recognized installed packages and provided type hinting and docs, but when I switched to asdf it did not seem to work. But, I was new to asdf and perhaps was using it incorrectly.
I was always assuming that when I'm in the command line, running asdf to use the right python works because the path is correctly established. But, when I run zed, it launches without the path setup step, and things went badly. I'm just speculating, but I could not get type hinting and didn't know how to fix it.
> I was always assuming that when I'm in the command line, running asdf to use the right python works because the path is correctly established
In a typical setup asdf is installed by sourcing some shell script in your .bashrc (which will then add shims to your path). It might very well be that Zed didn't execute `python` in an interactive shell, so the shims weren't available. There are various solutions here but the easiest is probably to add the shims to your PATH yourself.
As for venvs, using asdf doesn't mean you should no longer use venvs since all projects using the same Python version (managed by asdf) will still share the same site-packages folder. In other words: I'd still recommend setting up a venv, e.g. through Poetry or uv/Rye. Besides, once .venv/bin/python symlinks the asdf Python shim, Zed might have an easier time finding the right binary, too.
For e.g. I write “library” of which v1 depends on somelib.so.1.0.0 and v2 depends on somelib.so.2.0.0
If somelib has some symbols clashing in the names this can cause real problems!
There were issues relatively recently with -ffast-math binary wheels in Python packages, as some versions of gcc generates a global constructor with that option that messes with the floating point environment, affecting the whole process regardless of symbol namespaces. It's mostly just an insanity of this option and gcc behavior though.
Maintaining fine grained symbol versioning is a pain and a massive amount of work for the package maintainer:
https://invisible-island.net/ncurses/ncurses-mapsyms.html
Honestly, multiple installed versions like jinja1 and jinja2 sounds best to me.
Having your software depend on two different versions of a library is just asking for more pain.
BTW, I still need to fix it to run on 3.12+ in a neat way. For now, it runs, but I don't like it.
Though in principle, my preferred approach here would be similar. Manually install in a particular prefix, add it to the path (manually at runtime or programmatically via the sys module), and then import multiple versions under different namespaces...
I won't start the same thread here (though thank you for pointing that one out), but when I've come across scenarios similar to those in that thread in my own projects (which, admittedly, while not trivial, were probably not enterprise-scale), the solution generally still involved making sure that one ensured the paths / global variables were suitably modified in the relevant modules before the relevant calls, to ensure you're using the correct namespaces expected by the recipient. Which may be tedious, but to me is not absurd; you're just sticking to the contract expected by the recipient. The bigger issue here (for me) is probably whether those contracts are visible/explicit or not, and how/whether the contract is enforced in the recipient library ... but I would hesitate to call this a multiversion / dependency problem.
I haven’t used UV, but it says that it manages python as well as packages - I’m guessing like conda, python-venv, and of course nix does.
If the C api is an issue, it sounds like you have control over it if you need it. You manage the python distribution, so could it be patched?
This way it feels like you’d be able to establish not just what is being imported, but what is importing it - then redirect through a package router and grab the one you want.
This may be particularly useful if you’re loading in shared libraries, because that is already a dumpster fire in python, and I imagine loading in different versions of the same thing would be quite awkward as-is.
For lack of a better word, the single package version forces the ecosystem to keep "moving": if you want your package to be continued to be used, you better make sure it works with reasonably recent other packages from the ecosystem.
Semvers does not matter in this way. The issue with having a singular resolution — semver or not — is that you can only move your entire dependency tree at once. If you have a very core library then you are locked in unless you can move the entire ecosystem up which is incredibly hard.
And a very real issue is that young developers don't know anymore how to develop by limiting dependencies to the strict minimum. You have some projects with hundreds of dependencies without a real reason than lazyness or always using the new shiny thing.
If you are still struggling with this in 2024, you are missing the actual challenges of the world.
Imagine 2050, flying cars, talking robots, teleportation, and John Developer is going to be releasing solution 57 to a problem that was solved in the 90s
It seems the more users a language has; the more dev tools get written for it.
d-pad(pad_character,direction,string)
Of course this would be an internal dependency of both left-pad and right-pad.
Sadly what happens now is that everyone under the sun tries to evangelise and "create content" for these new tools so much that the natural filter mechanisms don't work. Doubly-so because tools like Google effectively created a new fitness function for peoples' behavior that incentivizes just plain old content creation (whatever weird form it may take, including new libraries being created + promoted).
Beyond just tool names, it's also important to realise that there has been a significant movement from the Python development team to standardise aspects of tooling. Tools like Poetry and uv weren't possible a few years ago before there was pyproject.toml to unify a bunch of separate things, for example.
At any rate I take care of all of my python installs by not downloading a gazillion of random packages online, if I ever reach the situation where package 517 depends on package 208 and package 598 depends on a different version of package 208, I'll just pull out the FlemmenWerfer and trash the whole thing before it reproduces.