The problem with packaging in Python
blog.ionelmc.ro
blog.ionelmc.ro
This is a really terrible suggestion.
The last thing we need is another way to do things, that only takes care of, say, 70% of the functionality.
All that will happen is, someone now has to look through the simple interface, and decide oh wait, I need one feature this doesn't provide. Now your "solution" has made the problem worse, not better.
If you're going to provide this sort of simplified interface, you need to make damn sure a super majority of users never need to look past it. Otherwise you've done much more harm than good, by providing yet another competing alternative.
The real problem with packaging in Python is that users are asked to write this program, setup.py, that by itself should not be a very difficult program to write. Its tasks simply are not that complicated.
What makes it complicated is that they're then told that their program should consist of a single function call --- a call to this monstrously complicated setup() function, with a confusion of conflicting arguments.
This is a stupid way to write a program! This design decision is at the heart of the whole problem. The interface to setuptools, distribute, distutils etc is fucked and always will be fucked, because there's no way to provide a good interface to the functionality via a single function call.
That's why every semi-complicated setup script ends up having to monkey patch the internals of setuptools or distutils, to say, replace some of the Extension class machinery, or try to correct a compiler flag. It's because the design is fundamentally a failure. The whole idea does not work.
Things would be much easier if we had direct, simple control. You can provide a default fall-through, but it should be easy to route the "build_ext", "install", "sdist" etc commands to your own function.
Just tell us what our program needs to do and provide us a nice library of building blocks to do it, and we will write this program. Simple.
By far the majority of my tickets for my NLP library, spaCy, are related to packaging and installation. Why should putting some files onto a user's computer and zipping up some files on my machine and calling out to a compiler and specifying some dependencies be harder than understanding natural language?
When a framework is done badly, in a tightly coupled fashion, it may make a bunch of basic stuff easy but it will also make more complex stuff either impossible or possible but requiring horrendous contortions in the code.
I think that a bunch of libraries is better than a poorly designed framework but I still don't think they're a substitute for a well designed framework.
Unfortunately, many framework developers' attitude to such contortions is not "wow, we screwed up, we'd better accomodate this valid use case" (or even better make it possible to do declaratively) but instead this:
"You shouldn't be doing that"
Now the opaque black-box "framework solution" works again!
Also discussed in my "API bondage" post: http://rare-technologies.com/data-streaming-in-python-genera...
I would say that once you've chosen the correct algorithm to implement and understood it deeply, writing algorithmic code often isn't so difficult.
In contrast, interfacing with existing systems can be arbitrarily complicated, if you're handed something with arbitrarily bad design.
Compared to that, I'd consider the overgrown setup.py interface or need for a MANIFEST.in file to be barely even a problem.
I tried writing a package that took care of specifying unix packages in a distro independent way (https://github.com/unixpackage/unixpackage), checked for installation and asked politely via sudo to install if the package wasn't installed.
However, in pip, the "ask politely" approach became impossible after v1.5.6 because of this bug/decision:
https://github.com/pypa/pip/issues/2732
(a sudo prompt came up but it was not possible to give an indication to the user of why it was there, so I gave up on the idea)
perl Makefile.PL
make
make test
make install
And package authors would use helper modules to handle the Makefile.PL cruft[0][1]. It sucked.rbjs wrote Dist::Zilla[2], a packaging abstraction tool: you would tell it what was in your package, and it would write the installer for you (and you could choose which of the Makefile.PL helpers to use, if you cared). The end user never sees any remnant of the tool.
The closest Python has right now is pbr[3].
[0] https://metacpan.org/pod/ExtUtils::MakeMaker [1] https://metacpan.org/pod/Module::Install [2] https://metacpan.org/pod/Dist::Zilla [3] http://docs.openstack.org/developer/pbr/
Python really needs one, and only one, true way to do packaging, but I think it's too late for that to happen now.
The overhang of backwards incompatibilities hangs around forever, too. Every time you invent a new and better way of doing it you have to continue supporting the old way for years. People will keep doing things the old way forever, too.
Python's package manager has many flaws I find that it generally works okay. I tried Go out a few months ago and once I saw that their approach to package management was to point your code at a github repository I backed away slowly and didn't look back.
The ongoing work to modernise stdlib and to incorporate the best parts of the various alternatives in a compatible way while balancing the sometimes conflicting needs of devs on windows, devs on *nix, sysadmins, distro packagers etc is very complex. It will take a long time to sort out in an evolutionary incremental way.
Part of the problem is that the 3rd party tools they are trying to improve/replace were never fully specified or understood. There are a lot of corner cases to find.
There is slow progress being made though.
Although it's becoming more and more common to require NPM as well for PHP development now, apparently.
It basically 'just worked' across Python 2, Python 3, Linux, OSX and Windows. There's a lot of outdated examples out there, and it might be more troublesome if I had a more complicated project, but... well, it was ok for me.
[1] http://learnpythonthehardway.org/book/ex46.html [2] https://hynek.me/articles/sharing-your-labor-of-love-pypi-qu...
Making .rpms and .debs is not all that hard, though I admit there are a few too many package formats (and that still leaves non-Linux users without a clean solution). Beside this, what benefits do users of these language-specific packaging systems enjoy that I am missing?
If they were better designed, you could simply tell debian users to install rpm to handle their rust packages, like you now tell them to install cargo. That sounds ridiculous, but it shouldn't.
I'm hoping Nix, Guix or something of that sort can one day fill that role. Integrate them nicely into distros and let them slowly take over the non-system packages. Let rpm and dpkg slowly wither away until finally, with a little "pop", they disappear into nothingness.
Of course, all these language-specific package managers support Windows, so I hope Nix and/or Guix actually do so. As git has shown, it doesn't even have to be amazing support. But, without something basic, people are going to keep inventing new package managers to make Windows work.
Will you kill off all other package managers? Obviously not, but I do think you could prevent quite a few people from re-inventing the wheel for their next language if there was already a good tool to build upon.
Personally I like node's way of using package.json and cargo's Cargo.toml. If someone made that for python that'd be awesome.
The other thing node has specifically is the "node_modules" folder, which recursively has all the dependencies and their node_modules folders. While it might not be super space efficient it's really stupid simple.
Also, packaging systems need to get better at integrating with VCS versioning. The elm packaging system goes beyond this to automatically bump versions when your api changes https://github.com/elm-lang/elm-package#version-rules, whoa cool. But at minimum, make a note of git/hg id.
The 'just use npm' approach is interesting too http://dominictarr.com/post/25516279897/why-you-should-never.... NPM's approach to 'dependencies of dependencies' has a lot of fans.
Finally: part of the problem here is that there's no package manager for C. The author dances around this in saying 'C extensions are hard, we should ignore them'. If we had a good C package manager a lot of work linux distros do in supporting C libraries would be simplified and builds of everything would be so much more portable.
NPM's approach works for Javascript because package A importing libfoo-1.0 and package B importing libfoo-2.3.11 doesn't cause any sort of global namespace collision for "Foo" within the JS runtime; each Foo is just a property of the object returned by require().
This isn't true of most other platforms; you can't e.g. import two versions of a Ruby library into the same Ruby runtime. (You can import two versions of an Erlang module into the same BEAM VM, but the second one will be treated as a code upgrade for the first. Processes that don't do a full-qualified tail-call will stay on the "old" version, though, so both versions can be loaded in parallel. There's still only one global "module table", though.)
On the other hand, the "packages get their own scratch namespaces which their dependencies can be exposed into" is true of Unix (or rather, can be made true by clever use of symlinks/chroots/etc.), which is why Nix can do what it does.
That said, it's not impossible to hack the import command to use a different package tree per module. The main difficulty would be supporting the different ways of storing modules (eggs, filesystem, .so support).
Everyone wants Nix.
But we're not allowed to have Nix.
Because the world sucks and people look at something that has a little bit of necessary complexity and think, "I don't like complexity! I'll build something simpler", and they build it, and it looks cool because it seems simple, but then you use it and it turns out it has implicit complexity.
I think in the end we want a self contained OS package, to offer a sane install method for any target OS. Why would users have to make a difference between e.g. a C program that they install through apt, or a Python program that all of a sudden requires a completely different way of installing?
Here are some interesting thoughts on it: https://hynek.me/articles/python-app-deployment-with-native-...
And here's a tool that I wrote to be able to wrap up a virtualenv into an OS package, while specifying a precise list of OS dependencies. It's not perfect by far, and admittedly a bit crude, but it already does the job well for some real life "production" stuff. It's called vdist (virtualenv distribution).
https://github.com/objectified/vdist
Documentation is here:
System engineers are charged with keeping various programs running on (usually) a fixed distribution/OS. They need to have a way to update SSL when there is a security issue, for all applications. They need to be able to update various libraries and have programs transparently use new or otherwise fixed versions of libraries. They also very much need the ability to uninstall packages.
OS level packages are a solution written to serve the need of sysadmins. But software engineers don't care, because maintaining the correct package list and versions for 4-5 distributions is horrible. So they don't.
What I don't get is why docker isn't the ideal solution for both groups. Assuming, of course, sysadmins have both ability and the inclination to insist on the sources.
For some things, Docker is a great solution for both groups, but that requires that your organization and application meet a list of requirements. Is the application only going to be run by you, as a server? If it's going to be run elsewhere, you may not have the option to require it run on Docker or if it's going to be run as a client then you definitely don't have the option.
The conflicting goals come from the System Engineers goal of making it run on all platforms vs the System Administrators goal of making it run on this platform. System Engineers want a solution that makes it easier to abstract away the platform, System Administrators want a solution to make it easier to install on this platform. OS packages are easier as a System Administrator because it all ties into a central source of authority. Each package knows which other packages are installed on the system and which packages it needs. Fragmenting the environment makes it more difficult because you've got distinct sources of authority. Pip package A needs version 1.2 of OS package B, but pip package C needed for pip package D needs version 1.3 of OS package B. All of which you'd have no idea about until it all crashes and burns and you have to go figure out why.
OS packages are usually better in terms of installation, but pip packages are easier to create. Ubuntu's new packaging system is meant to address this, though we'll see how it goes.
I've done similar for Ruby projects using many gems with various C library dependencies, though I used a fair bit of custom Nix definitions for this (i.e. at the time, the state of Ruby packaging in Nix was poor; I'm not sure of the current status).
A signifiant requirement is that it must live in /nix (it's theoretically possible to build everything for another path, but I don't think many people do this so there may be surprises).
And this is all easier if the target system has a Nix package manager, rather than dealing with (somewhat bulky, since includes everything) tarballs.
For me its much cleaner and controlled than distutils plus the knowledge is largely transferable to other stacks. I wrap our node, java, and ruby applications using the same tooling as well as any services we depend on. (kafka, kibana, spark, etc) After that I can pass off making images, containers, installing, managing stuff, whatever onto anybody in IT or operations that knows linux.
Obviously that wouldn't work well for everybody but it sure beats trying to manage 20 different ways to do everything in my case.
In the last couple of years setuptools changed / broke so many things that previously worked without documenting anything. Even trying to figure out what the relevant docs would be, and who maintains them is hard. As Nick Coghlan mentions in the comments: Even the people maintaining this stuff have no idea what it all does.
And still, the correct (but incredibly boring and frustrating) solution is for someone to sit down and figure this stuff out and document it. Definitely not me though. :)
"""Python Build Reasonableness"""
http://docs.openstack.org/developer/pbr/
It is pretty nice and helps a ton with packaging sanity...