Python Wheels Crosses 90%
pythonwheels.com
pythonwheels.com
I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
We still used wheel archives etc for managing some python dependencies during our dev and build process, but it was never a complete solution: some python packages depend upon non python shared libraries so you need to install them using an operating system level package manager. Then when shipping the software to users you cannot assume the user has a working version of python, or even if they do you don't want to increase your support burden by having your application running inside some arbitrary python environment.
E.G:
- shiv leave files hanging around, but you can address those files directly. With PyOxidize, you will need to use https://pyoxidizer.readthedocs.io/en/stable/config_api.html#... a way to read a non Python files. Because most projects don't know this and just use open(), they won't work out of the box and you may need to monkey patch open().
- pyoxidizer requires a compiler, which is easy on linux, but a higher requirement on windows. Compared to shiv, which is just a pip install away.
There is no perfect solution yet, because you have really often many things to balance: bringing in the python vm or not, allow compiled extensions or not, provide support for open() or not, etc.
Sure people may not use the tools they have but then again people also don't upgrade their OS packages or maven packages. So for those people packaging everything doesn't hurt any more since they've got rampant CVEs already.
Like if building a pet app is your company/team's job then containers more than fine. If you're a distro maintainer that has thousands on thousands of packages then sharing libs starts to make a lot more sense.
They're screwed. Hopefully you aren't.
Hopefully.
The problem is that distributions get their software from upstream developers. And if upstream decides to go static or use embedded code copies, the maintainers have much more work trying to untangle everything.
In my experience, packaging Go software for Linux distributions is very frustrating and tedious because of that.
Packaging Python or C/C++ software is much easier.
Wondering since the application developers I know seem to prefer static linking.
I can't help but wonder if we'd all just benefit from better build systems, so that "rebuilding all downstream consumers" wouldn't be hard.
Perhaps this "untangling" isn't the best idea, at least not always?
- the bundled things might have unclear licensing
- bundling stuff that is available in the repos is sub optimal, as the repo version will be updated for security issues while the bundled version will not
- you can use some mechanism to mark what stuff is bundled in a package, but then you need to make sure the bundled thing is patched and rebuild each thing bundling it, where with system wide dependency you just update that and you are done
I'm fairly sure that almost all projects are receptive to updates for dependencies with security issues; I don't see why a distro needs to "untangle" anything for them.
> with system wide dependency you just update that and you are done
Unless the system-wide update won't work with your program.
The entire point of the Go tooling is that you can reliably and consistently build your software, and replacing random parts to create a weird unsupported (by the actual developers anyway) version is something the Go tooling was never designed to support: it goes exactly against what it was designed to do. Hence my previous comment: perhaps this isn't the best of ideas.
- Outdated packages
- Having to build from source when the package is not in the repos
- Failed builds for no good reason
- Unability to have many versions of that same packages
- having to install log(n) build tools if I want to build n packages
- etc
while in half-developed OS'es everything just works (application-wise, at least)
There's a fundamental unresolvable tension in automatic update systems between "not updating breaks things" and "updating breaks things"; you cannot solve for both at once and either one will get users mad at you, although not upgrading at least gives a pushback argument. See how all this has played out with Windows and Mac updates, for example.
You really should check what APIs your dependencies provide and how stable that API is, then set minimal dependency version. If you set the dependency to a specific version, you are creating a headache for anyone trying to use your software later on (by potentially forcing them to use oudated and insecure software) and for distro maintainers trying to keep distro software up to date and secure.
What if the dependency is so unstable that you have to pin the version or even use custom patched version? Well, thenaybe its not somethong you should be depending on & rather use something with less features but more stable. Othervise you are really being unresponsible - using quick convenience at the cost of long term usability of your software.
The thing is that in general you don't want to have many version of libraries available in the distro in parallel as each has to be separately maintained, security patches applied, build issues fixed, etc., eating valuable maintainer time. Also in some cases libraries of different version can't coexist on a system cleanly.
Then imagine every piece of software just pins versions of their dependencies to a specific random version that happened to work at the time. To satisfy those arbitrary dependencies the distro would have to maintain all these versions at the same time, which is simply impossible resource vise, not to mention incredibly wasteful (both in maintainer time & system resources).
As for stuff breaking if you don't pin dependency version - well, distros have mechanisms to handle that. For example for Fedora, there is a stable release very 6 months & stable releases are not expected to get major changes in libraries, just bug fixes and smaller enhancements.
And at the same time there is a rolling version of Fedora called Rawhide, where all the latest package versions land and where integration issues are addressed. So any breakage would happen on Rawhide and be addressed by maintainers (of the library/software affected or both) long before a new stable release is cut from Rawhide and users will actually use it.
For an example I'm maintaining the PyOtherSide Qt 5/Python bindings on Fedora. A while ago the build failed due to Python being updated to 3.9. I reported the issue upstream, which quickly fixed it and I've built an updated version in Rawhide. All this long before a stable Fedora version will get Python 3.9, but I can be sure that when this happens, all will work fine.
[0] Remember when Debian generated predictable random numbers because a maintainer wanted valgrind to shut up?
Now in comparison people using NPM just blindly pulled random stuff directly from upstream without anyone doing any sanity checking at all - no wonder one package vanishing made the whole thing fall over, often directly in production.
Especially with Python, and especially with the 2/3 split, people got used to assuming that the distro version was something broken to work around (e.g. Redhat, OSX), and that all "real" work happened in one's local language-specific package cache or venv.
I'm increasingly of the opinion that it's a mistake for distros to ship Python or Ruby packages in their distro-specific package format, but I can also see that's going to be a holy war.
As Linux user who sometimes build packages (tinker around) I need all kind of dependencies - python, ruby, perl, haskell. It is much safer and faster to use distro packages.
Breaking changes on major version is awful for any consumer. Python 2/3 story is a shame. These can't be arguments against distro packages.
I can't help but wonder if there's a better way that could scale to open source. I personally like monorepo-based development a lot - where you have one version of every library for the whole repo, and a total ordering on changesets (and a strong test suite to catch regressions). But organizing open source into a monorepo or even a virtual-monorepo seems tricky. Might be the sort of practice that can work well for companies but not so much when decision power is more distributed.
And users would find outdated versions (maybe unmaintained). Current system update dependencies automatically or allow collaboration on life support (patched application, dependencies, config). It would be good to have alternative like container or virtual machine with snapshot from the days long past. But it should be clear this is not safe.
Em, like all other packages?
Here are the scenarios:
- dynamic responsive remove bug: Positive/neutral. Team X would have done it anyway.
- dynamic unresponsive remove bug: Positive.
- dynamic responsive add bug: Negative. Team X will see the bug but only be able to passively warn users not to use Y version whatever.
- dynamic unresponsive add bug: Negative. Users will be impacted and have to get Y to fix the error.
- static responsive remove bug: Positive/neutral: Team X will incorporate the change from Y, although possibly somewhat slower (but safer).
- static unresponsive remove bug: Negative. Users will have to fork X or goad them into incorportating the fix.
- static responsive add bug: Positive. Users will not get the bad version of Y.
- static unresponsive add bug: Positive. Users will not get the bad version of Y.
Overall, dynamic is positive 1, neutral 1, negative 2, and static is positive 2, neutral 1, negative 1. Unless you can rule out Y adding bugs, static makes more sense. Dynamic is best if "unresponsive remove bug" is likely, but if X is unresponsive, maybe you should just leave X anyway.
If the install is on a consumer machine for regular usage then the right answer is shared libraries for the machine. It adds a lot of complexity for the packagers for that OS but you get a lot of safety for that consumer.
If the install is an inhouse app deploying to servers a company controls/rents whatever then the right answer is probably a static application. Consumers are more likely to want to chase the latest version of everything. Company developed software is far less likely to want that. The problem then becomes less about packaging and more about deploying fixes quickly. A static binary that is rebuilt, tested, and deployed is going to be a smoother path to fixing that error than figuring out how to deploy your new shared library and avoiding any issues in our production environment caused by clashes to that library. This is made more problematic by the likelihood that you will need multiple versions of that shared library with the fix to meet the needs of applications that can not yet be upgraded. Static binaries make things easier in this scenario and thus quicker to resolve.
Most python packages I have seen will put their python source in the python site-packages directory. Where were the source files you were trying to load located on the file system?
> "static" binaries that package everything together into a single, self-contained unit.
Ummm, except debugging "static" binaries without symbol file is hell.
If there's a .py source somewhere on the disk, I can at least try edit the source file from the package and debug internal private variables as last resort.
> Trying to load source files from all over the filesystem at runtime is hell.
This sounds like a mess unrelated to Python. Why are you trying to load source files from "all over the filesystem" at runtime?
> I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
Compiling source files from all over the filesystem would be equally annoying.
It's basically a huge mess. There's even an XKCD about it.
Consider the alternative: Go compiles programs to a single statically linked executable with no dependencies. You literally just copy one file to your target machine and run it. It basically can't fail. That's part of the reason Go is so popular for server stuff.
And it's not just because Python isn't compiled. Other scripting languages handle this much better. Even JavaScript - for all the hate node_modules gets for being enormous, at least it works reliably!
> Those libraries are installed by Pip... somewhere on your system.
They go into your virtualenv.
> And then Python has to find it somehow at runtime.
If you use a virtualenv, this works 99.9% of the time.
> That relies on environment variables being set correctly
What environment variables? Just use `./venv/bin/python` directly without environment variables.
> a huge mess even before you consider things like virtualenv
There's nothing to consider, always use a virtualenv. It's the same for thing for Node, except it implicitly handles it for you.
> and the fact that non-Linux systems usually have multiple copies of Python installed (and Python 2 and 3!).
How is that a problem? Just pick the interpreter you want to use when creating the virtualenv:
virtualenv -p /usr/bin/python2 venv
virtualenv -p python3 venv
virtualenv -p python3.7 venv
> Consider the alternative: Go compiles programs to a single statically linked executable with no dependencies. You literally just copy one file to your target machine and run it. It basically can't fail. That's part of the reason Go is so popular for server stuff.I agree that Go got it mostly right and it just works, except for that fact that it didn't even have a package manager for like 10 years so pinning dependencies was impossible unless you forked repositories.
> Even JavaScript - for all the hate node_modules gets for being enormous, at least it works reliably!
In what way is Python + virtualenv + requirements.txt less reliable?
> It's the same for thing for Node, except it implicitly handles it for you.
Indeed.
Funnily enough node_modules is one of the main regrets Ryan Dahl, the creator of Node.js has: https://www.youtube.com/watch?v=M3BM9TB-8yA&t=755s
Isn't that why Docker is so popular? If we had Python static binaries (or the ability to have several versions of the same Python library installed) would we use Docker so much?
* insert lines of code into asciidoctor Ruby to debug how the command line interface differs from the library interface;
* read through Raspberry Pi SenseHat code to figure out that the drivers won’t load via SSH (you have to plug in an HDMI cable!);
* rummage through Python’s PY Sequence List to answer the question “what kind of sort does Python actually use?”.
The problem with shipping binaries is it lowers the standard for being able to build from source.
If a package is easier to ship as a binary because the source is hard to distribute reliably, we give up hope of end users being able to modify and build their own versions.
This is obviously bad for freedom. It’s also hard for debugging and learning.
I always _try_ to show how one can answer problems from first principles before resorting to “go look it up on stack overflow”.
It boosts ones sense of usefulness as a teacher if one can occasionally give more in depth answers as teachable moments.
For one thing, demonstrating deeper understand is how you score high marks in exams.
Open up Python's source code and check - this is first principles-first hand.
As a matter of fact I am going to go as far as saying that reading source code is the _only_ way to actually be any good in software engineering because it's the only source of truth that never gets out of sync with the de-facto.
Did you expect to find a blog post about it (you can find it on the first page of Google search I mentioned)? Or a YouTube video (didn't pay attention to that)? Books are reasonably good but usually get out of sync quite fast.
And going back to first principles: the reason most of the literature may not mention what sorting algorithm Python uses is because a) it's irrelevant as a first principle knowledge b) it gets quickly out of sync.
An important part of school is learning how to learn. It is a constant struggle with pupils to get them to actually think instead of googling for an answer.
Sometimes it feels a bit like I am a grumpy oldie telling them things like _in my day we had to look things up in books instead of using Google!_ but most of the time it’s valuable in and of itself to go through things from first principles.
(The sensehat bug was a case in point.)
As far as source code is concerned, it turned out to be quite tractable to follow the Python source code at least far enough to find the timsort.txt document.
If any of these kids go on to have technical careers, there will be plenty enough time for googling answers later on in life. Hopefully some of them will be writing the answers too.
No one in tech support believed me. No one was willing to help me. They told me to make a new account.
I've done my best not to buy from Amazon since.
Please don't be hostile to users.
Possibly after you provide some kind of validation like the last few digits of credit card tied to that account or something?
I'm really curious about this. Why would a driver care about whether you load it on a "real" tty vs. over the network? How does it even know?
We didn’t dig further than that, though the error was quite obscure until we read the source code, at which point we realized that the “FB” the error mentioned meant the Linux console framebuffer.
I assume the Python module for operating this device ships with a GUI, which is couple to some Linux framebuffer initialization code, which in turn only works if the framebuffer is enabled, which presumably it isn’t if the host boots without an HDMI device connected.
The LED matrix is an RGB565 framebuffer with the id "RPi-Sense FB". The appropriate device node can be written to as a standard file or mmap-ed.
https://www.raspberrypi.org/documentation/hardware/sense-hat...
You might be giving up hope too soon, because what you describe has in my experience been the life story of source code from the start, to some extent. And I'm not blaming it for that. Distributing source code (well, not the distribution itself, but getting it to build) on all possible systems out there is hard and impossible to get 100% correct. It's simply inevitable to encounter systems on which it doesn't build because of their specific configuration, or user error.
Despite that, end users have been building their own versions and still do. Users who really want to don't let themselves hold back by a compiler error. Granted it's possibly that in a perfect world where everything would just work there would be somewhat more users doing that, but there's no such thing as a perfect world.
Moreover I do have the impression things are actually much better these days, mainly thanks to CI which makes it easy to get out of the bubble and build on multiple platforms. At least I don't have that feeling anymore of 2 decades ago where I sort of was prepared to waste hours getting anything to build which got handed to me on the internet. Still, it would be very interesting to go and see the percentage of projects out there which just build according to the instructions.
What? I can assure you this is not true.
My R Pi has never touched an HDMI cable and runs a sense HAT just fine. What OS?
If you needed to dig through source code to find out that python use Timsort, I genuinely am curious about how bad you are at search engines.
I believe in reading source code to answer some things, but that example is really odd for anyone who has researched sorting algorithms since Python literally invented a sort used by many other languages now.
If you need a billion dollar search engine and an internet connection to make sense of the source code on your local machine, I am genuinely curious how bad you are at reading source code!
But as I said, that’s pretty uncharitable. Because of course I didn’t have the source code for Python locally.
I downloaded it after Googling “Python source code”.
+ Explaining something that is "obviously common knowledge".
+ Reducing motivation of participant to share their experience.
I think we have HN bingo right here.
It isn't always possible, but Nix places a lot of emphasis on building software from source (or downloading a pre-built copy from a binary cache), and it has fairly good support for overriding/customizing the expression that builds one or more pieces of software.
A few months ago I was trying to debug a test in the Oil shell project that broke only on macOS for very unclear reasons. I eventually figured out that, in a specific mode, bash (which both Nix and macOS use as sh by default) was filtering out a specific environment variable that this specific test was trying to set, before running a target program with a sh shebang. I only needed to add a few lines to the project's build expression to patch bash and use the patched version for the tests (all without affecting the system/user bash).
The key benefit that nixpkgs has over other Linux distributions' package sets on this topic are:
1. It allows you to _program_ your package set, thus adding static-linking overrides to 1000s of libraries/executables in a go, instead of having to modify each package manually.
2. It allows the end user to choose what linking they want, as opposed to the Linux distribution making that decision for them upfront.
3. It allows to do this from any Linux distribution.
In my experience, the best way to read source code and experiment are to get the source from the upstream repo (eg. GitHub) and read/build from that. I do this with code written in C or C++ all the time.
The code is always out there (unless you're talking about proprietary software, which none of your examples above are).
I've suspected that for a long time without confirming it personally, so thank you for your time.
And yesterday I spent at least an hour fighting with cargo build. Without knowing exactly why we fought both our anecdotes are quite useless.
You mean Docker containers? :-)
For me, most of the pain with the Python packaging went away after I started using Pip-tools[0]. It's just a simple utility to add lockfile capabilities to Pip. Nothing new to learn, no new filosophies or paradigm's. No PEP waiting to be adopted by everyone. Just good old requirements.txt + Pip.
Python has the most unpythonic package experience ever. As elegant as the language is, packaging is a convoluted nightmare. It still might work for some project, but certainly is no fit for the average python user.
That being said, I'll admit I don't know much about what it means to be pythonic. What do you think the drawbacks are to pipenv-type approach and what would be more pythonic in your mind ?
I wish I could just as easily say: ahh this script I made should become a package that I can reuse on my server. But then the server has Python 3.6 while your script has 3.8, so you end up installing another python, painstakingly watch not to install over the old one etc. When installing modules you have to do the same or set up a venv/pipenv.
All along the road there are possible ways to shoot yourself into the leg. Meanwhile in Rust you do a cargo new foobar, work on it, add dependecies to the cargo.toml, copy it to the server, build a binary with cargo build --release and copy the binary into the PATH. You spent zero brainpower on not breaking things and had capacity to think about other stuff.
The closest thing we have in those terms in pythonland is poetry, which is a very good start. But this should be part of the language, not something one has to do extra.
Thanks for the insight and clarification.
IMHO python, node.js, ruby, and other scripting languages share a historical confusion of concerns related to development tooling and attempts to counter the unfortunate choice of installing dependencies globally that tend to leak to the server side in inappropriate ways. With Docker.
There are other things docker can do (like mapping volumes) but I would be much happier if we solved it at the level of the executable.
I wonder if there have been any surveys on poetry, and whether it's trajectory (presumably exponential) suggests that it's going to become a dominant package management utility in the next 10 years or so?
https://stackoverflow.com/questions/3430400/linux-static-lin...
One need not depend on the filesystem for such things. And as bad as python packaging is, others get it better. One can have less-than-static-binaries without being quite that bad.
> I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
Plenty of problems with that too!
tldr; Do not use Python. It makes certain tasks seem easy, but all it truly does is make you look like an idiot developer. Pick literally any other language, as Python has screwed up too many basic concepts (related to POSIX and expectations from C) to be taken seriously.
I stand by my opinion that Python is not well-suited to applications written by a professional entity, intended to be shipped to 3rd-party clients. Python plays too loose with the fundamentals provided by the operating system to be a professional language.
The other issue is that there's no equivalent of the %(config) RPM spec directive to prevent the config file from being overwritten if it already exists on the file system.
So, for libraries, wheels are a good cross-platform packaging solution, but not so much for applications that require configuration files and init scripts.
I don't think the Wheel specification was intended to replace application distribution. It is aimed at the distribution of individual Python libraries.
Wheel is an unpack-into-any-prefix (and stays in that prefix) binary distribution format.
Libs are for the dev: you want the source code, you want to manage versions, you want it to integrate with your dev env, you want a well delimited scope, a way to install it using your usual setup (pip/poetry/etc) and to isolate it from project to project.
The established behavior of python setuptools when creating an RPM spec file via the bdist_rpm command was to create a spec file that would invoke the build command in the %build section and the install command in the %install section. The install command would place those files in the absolute paths specified in the data_files named parameter.
It has worked this way even back when distutils.core, the predecessor of setuptools, was used.
As for whether it's desirable, if you're missing the config file, init script, and other things necessary for an application to work, then you essentially have an incomplete installation and will have to copy those files from other sources and update them yourself via config management before you're able to run it.
> I guess I'm just asking, aren't these two separate problems (config management and Python packaging)?
In my view, config management should update a config file that's already there and also restart the service as needed (either when the package is updated, the config file is updated, or both). It shouldn't be responsible for placing the config file and init script itself.
For example, if I install an application like redis or rabbitmq-server, the package manager will place the included sample config file in the correct directory (e.g., /etc). I can further modify that file as necessary before starting the service. It will also include the scripts necessary to start the service as a daemon. Config management could be used to update the config file for environment specific concerns, but it doesn't need to essentially duplicate what the package manager has already done (create the directories, set permissions, place the file, etc). Python applications, regardless of how they're packaged, shouldn't really be different in this regard.
You create a RPM _from_ a Python package. The RPM (and associated software) writes init files to the correct place. This bears no relevance on wheels, which absolutely should not be filled with district-specific knowledge about init scripts and should definitely not write to random hard-coded paths.
Where is that documented? The fact that setuptools and distutils.core supported building RPM packages from a setup.py file indicates that python intended to support application distribution. Why remove that capability when moving to wheels?
> This bears no relevance on wheels, which absolutely should not be filled with district-specific knowledge about init scripts and should definitely not write to random hard-coded paths.
One thing I noted while figuring out this issue is that running:
pip install git+ssh://git@github.com/org/project.git@1.2.3
versus
pip install project==1.2.3
where project was packaged as a wheel on pypi had different results with regards to where files referenced by the data_files named parameter in the setup call were placed on the file system. Given that, was the change intentional, or is it a bug with how wheels handle package installation?
How is this supposed to work in virtualenv environments? What happens if you have two of them?
Part of the motivation for using wheels in our case was to allow for pulling in later versions of certain python packages as part of the dependency resolution process when installing or updating the package. Using RPM while attempting to get the updated dependencies would mean having to repackage every single dependency (direct and indirect) from pypi as RPMs and upload them to our yum repo so that yum would be able to handle dependency resolution.
> I don't think anyone advertised them as such. They replace eggs or installing from source .py files,
The setup method from setuptools does have the necessary features for handling system-level files and the bdist_rpm command leveraged that when building RPMs. It just seems that if this was a bug rather than a feature, then it should have been noted when transitioning from egg based distributions to wheels. Nothing in the documentation for eggs or wheels states that they don't support deploying system files in absolute file paths.
The fact that doing something like:
pip install git+ssh://git@github.com/org/project.git@1.2.3
and
pip install project==1.2.3
have different results for where files listed in the data_files named parameter of the setup end up on the file system suggest that the lack of absolute path support when installing wheels is a bug rather than an intentional change.
If everybody knew that, I think the confusion would mostly disappear.
So I was reading this comment and thinking to myself, how does one even pronounce "ai"? Language is not my strong suit. After searching the web, I found out that "ai" is a diphthong and it sounds like "eye" or the letter "i". Leaving that info here for others who may be confused. Sound: https://www.youtube.com/watch?v=uyKgPH0kmrU
I’ll keep this in mind to next time I want to write out the letter I and use eye instead.
PyPy: “pie-pie”
Of course Python imports can have side-effects, so you can achieve the same results with 'import malicious_package', but the installation avenue was surprising to me at the time so I created a simple demo. Also consider that 'import malicious_package' is typically not run as root whereas 'pip install' is often run with 'sudo'.
Never, ever, do a pip install in your non virtualenv environment.
Like, I understand it's great for managing package dependencies/setting up environments etc for a single project. But I've found it's an absolute nightmare for building docker images/generally doing build stuff when it's involved.
Not to mention conda installs seem to take waaaay longer than pip.
apt + pip works far more sensibly and reliably in my experience.
Pip, on the other hand, often "wins" by not even bothering to notice the versions of low-level libraries, etc. If that doesn't work, you get to keep both pieces.
[1] https://github.com/pypa/pip/issues/988
[2] https://pyfound.blogspot.com/2019/11/seeking-developers-for-...
A big thing for me is that apt + pip means I know more about what I'm installing, whereas conda seems like it'll go off and do what it thinks is best. If it's going into production then I want to know why we need package X and how it's been installed, rather than "conda says we need it".
Basically, I think conda encourages "install and forget" which means people don't really know or understand what they're putting into production. And that can cause a lot of problems further down the line.
Also the fact that conda installs everything to sandbox causes it's own issues. Suddenly I have two versions of a system package. Now I've got to do extra work to deal with that.
Then again, it could just be I never got over the time where uninstalling anaconda ripped apart the python install on my MacBook. That's a weekend I won't get back.
__
Also, Production shouldn't use virtual environments in my opinion. That's an additional deployment/build step which could fail one day. The container image is a virtual environment in and of itself anyway!
Re the extra "virtual environment", you can just use conda's base environment in prod. Here's how to do that: https://haveagreatdata.com/posts/step-by-step-docker-image-f...
But I often need to install via system package management for other dependencies. conda doesn't respect the base system package manager. That is what causes headaches.
If conda respected system package management first, then installed as necessary, I wouldn't have a problem with it as an admin. But it doesn't because it's not built for engineering/admins (want stability + efficiency), it's built for scientific projects (want to run code easily).
Also, I'm using the "royal* we. Like, we as in admins generally. I'm the only admin in my team (voluntarily), so I need to be ruthless with this type of stuff.
EDIT:
I think you missed my point about virtual environments.
The entire container is a virtual environment. Why would we want to use another virtual environement for no reason except the fact that conda wants us to?
It adds extra steps which we'd have maintain. Which means more developer resource spent on maintenance. Which means less time spent on new features.
It's just another thing that could go wrong. Simpler systems break less often.
My use case is this: as a data scientist, I start new code bases all the time. Each project, simple experiment, data analysis, etc. needs its own cleanly separated dependency environment so I don't end up in dependency hell (I have 12 conda environments on my machine right now). Conda allows me to handle these environments with ease (one tool and a handful of commands -> as detailed in the article). With conda, I also have my data science Python cleanly separated from my system Python.
Of course there are other tools that can handle this use case. But pip alone won't do the trick. I don't like to have three separate tools for this (pip + venv + pyenv).
When I put something into production, I naturally want to keep using my conda environment.yml and have the same environment in dev and prod instead of switching to pip + requirements.txt, which might introduce inconsistencies.
It's a nightmare of interop. Yes, it works for one person on a single laptop, but my experiences with conda outside of the happy path are universally terrible.
I'll stick with the standards.
Last time I could only solve it by deinstalling everything and reinstalling an updated version. Took ages to get all packages running again. Luckily my code still works but now I get depreciation warnings (and not in my code). Who knows what will happen next time. It only took like a day I needed for actually working with the data as well... urgh.
It's great for that purpose, but - as far as I know - it has not much to do with general Python dependency management.
Note that you can use 'pip' within a 'conda' environment as well, so you can sort of have it both ways.
(Disclaimer: I am one of the maintainer)
Edit: A web search points to 2012, so maybe it's "only" 8 years?
Edit 2: Pip came in 2008, so something changed somewhat before 2010 as I remembered. But what did it install if not wheels?
Before 2010? That's what I believed to remember. But it looks like wheels did not exist before 2012, so that can't be true.
and why companies like travis / github aren't more active in language-level packaging work
github gives away so much free docker time -- faster installation would save them money directly
Even more so with pep517 and being able to use different build systems.
Maybe it's not reliable enough?
In any case, I can see why PyPI wouldn't want to increase the scope of their work. Sometimes is super valuable to just be good at one thing (distribute what others give you).
Wheels can be generated out of thin air, there is no requirement that setuptools or any tool is ever involved, if you wanted you could zip up a particular directory structure and it is a wheel.
As for setuptools:
There's a class that setuptools relies on to do the actual building of the Python project, called distutils.dist.Distribution which has a function named is_pure():
https://github.com/python/cpython/blob/e42b705188271da108de4...
Which makes sure that it doesn't have any known external modules or c libraries.
So to do detection of whether an sdist is a pure python, you'd have to execute setuptools and setup.py because of course setup.py isn't just configuration, it's full on Python.
So just looking at an sdist statically you can not know if the sdist is a pure python project or not.
pypi could probably generate a lot of wheels. In fact, `pip` generates wheels from packages it downloades before installing them, so wheels can be generated for a lot of packages already directly from sdist.
Why it isn't done already? Good question. I guess it has to do with a majority of python packages installing just fine from source (those that aren't a problem for packaging anyway), and then those packages that need wheels because they use native extensions kind of are high-maintenance to bundle and roll out, especially if upstream maintainers haven't invested in building their projects for that.
I would love to see auto-bundled wheels. but I guess for this to fly, package maintainers would need to start standardizing a bit better (that also involves tooling). Easy things are missing, for example in setuptools you cannot even define smoke tests that distributors could use to verify successful bundling.
I use Python extensively but the "Advantages of wheels" section on this site is way over my head.
edit: Thanks everyone :)
Is anything in that unclear?
In practice, this bit I think it's pretty important:
> Wheel has a richer file naming convention. A single wheel archive can indicate its compatibility with a number of Python language versions and implementations, ABIs, and system architectures.
https://packaging.python.org/discussions/wheel-vs-egg/
Having packaged wheels but not eggs, my understanding is somewhat limited, but eggs being unversioned and having weaker filename conventions sounds like it would be a pain.
I mainly use binary wheels for projects that include compiled C extensions, so I need a different wheel for each supported platform. There's a straightforward filename convention for that, and all the tools support it, so I don't have to even think about that when generating or uploading wheels.
Eggs contain compiled Python bytecode which requires publishing egg versions for every version of Python. While Wheels are distribution formats that can support more than a single Python version.
* Do Ruby gems contain all binary depedencies statically linked into them?
* Is there a list of exceptions that are guaranteed to be present on target platform (see manylinux [1]).
The fact that you answered "Gems" rather than "something better" means the answer to both must be "yes" but I wanted to check.
Gems are more similar to Python sdist’s. Native code is compiled on the installing host.
There is no specification like manylinux1,2010,2014 like there are for Python Wheels.
In this sense, Python Wheels can be considered more sophisticated than Ruby Gems.