Making pip installs a little less slow
pythonspeed.com
pythonspeed.com
I'm in a disprivileged location, and it seems pip can only download from pythonhosted.org at a rate of 10-20 kB/s. Worse still, pip downloads timeout and fail extremely quickly.
If I rerun the pip install, instead of resuming the download, it will download the file from the beginning, then timeout and fail again somewhere in the middle.
I tried going through a private VPN hosted on Linode, with similar results.
For the official Python package manager, this is simply unreliable and unacceptable behavior.
You pip install from it instead of pypi. It will in turn download from pypi and give the result to you. It will also keep itself updated, and you can batch download during the nigh packages you assume you will work.
As a result, our entire team always as most packages locally available. Changing machine, location or purging cache didn't mean loosing this benefit.
Besides, pip caching wheels doesn't mean it's not making any requests, so it's still a better experience.
For big companies, I recommend it anyway: it speeds up the entire team work, CI, allow you to publish private packages, etc.
If you can't do that, the next best things is to do "pip download", instead of "pip install", and save the wheels into a hard drive.
EDIT: And thanks for the tip from godmode2019. I upvoted you both.
However, they don't change the fact that pip by default assumes a decent Internet connection, with its short timeout and no resume on downloads, and thus is unusable with anything less.
Again, IMHO that's simply unacceptable for the official Python package manager.
I would say that Python, and especially pip, is actually relatively friendly to scenarios where non-obstructed Internet is not generally available, compared to many other offerings from other programming languages (especially those without corporate backing). Devpi was already mentioned as a solution; in fact, since pip’s --index-url accepts a file:// URL, you can even simply pre-download the wheels, arrange them in the correct hierarchy (see PEP 503) inside a thumb drive, and just pass that around. The --find-links --no-index combination may also be an interesting approach for simple setups. There are a lot of things to try, before you conclude things do not work.
1. AFAICT pythonhosted.org is hosted at fastly.net, which seems to be heavily throttling downloads from certain disprivileged locations, as well as downloads from cloud providers such as Linode. It would help a lot if they could ease up a little on the aggressive throttling.
2. Make the pip download timeout longer to better accommodate spotty connections.
3. Make pip downloads resumable so that a download makes progress each time pip install is run.
4. Make pip download each file over multiple HTTP connections in parallel. Download throttling applies to a single connection, so downloading over multiple connections will speedup the download.
Thanks a lot for listening!
Seemed like a great idea, and perhaps I just needed something beefier than an RPi 3, but it didn't work out for me.
For example, as a Chinese, I often switch to `https://pypi.tuna.tsinghua.edu.cn/simple`. You may need to look up which one is available and faster in your region.
But I shouldn't need to know this, pip should have taken care of picking the best download mirror.
Or it could just support resuming a previous download. I don't actually mind waiting for a slow download or having to rerun pip install a few times, as long as it makes progress each time I run it.
Some discussion: https://discuss.python.org/t/why-does-pip-reach-out-to-indic...
I don't recall if this issue is solved by specifying hashes (or perhaps even pinned versions are adequate) -- I would hope so.
> adding additional mirrors (instead of switching mirrors as suggested here)
Pip install git+<repo_url>
pip-compile # resolve and pin the dependency tree
grep '==' requirements-pinned.txt | xargs -n1 pip wheel # build every wheel in parallel
pip install *.whl # install the pre-built wheels
It may not avoid all the slowness, but when nicely integrated into a build/CI system it can avoid suffering it more than once.It still happens from time to time, but if somebody else read this comment, you want to check how many wheels you build before doing so.
Also, you may want to cache those build once and for all.
A lot of maintainers don't pay attention to excluding test cases and artifacts from their packages, leading to ridiculous package size growth.
Last I checked numpy was close to 100MB, a huge chunk of that being test case artifacts.
Django packages all translations for all languages to ever exist.
Pandas bundles pre-built dynamic libraries with debug symbols.
Etc.
Most of my "production" virtualenvs are close to a GB nowadays, which is insane.
I did have a look at numpy and on my machine the tests did not bloat it as much as you made me believe. The `core/tests/` modules are 3MB and the `__pycache__` doubled it to 6MB. What do you refer to as "test case artifacts"? The modules, pycache, or both?
Also, wrt you statement on pandas; is it the debug symbols that account for the bloat, or the libs themselves?
Just installed a quick "data science" like virtualenv, here is the example things I talk about:
$ du -sh scipy 110M scipy $ find scipy -type f -name '.so' -exec strip -s {} \; $ du -sh scipy 76M scipy
$ du -sh django/contrib/admin/locale/ 5.4M django/contrib/admin/locale/
> Sure, but it's straightforward to install those [0]. A separate test suite would not be.
0: Those = Other packages you would have to download to make use of the testing components, i.e. pytest
On a related note: our NPM install times are even worse: any tips to help there would be welcome.
The negative: sometimes it plays bad with some packages. Sometimes you need the (stupidly named) —shamefully-hoist, sometimes even that doesn’t help.
pnpm by default does NOT have the flat structure in node_modules but puts into separate sub-folders; sometimes it plays bad with some dependencies that expect the flat structure (usually webpack-related cludges)
But if you are starting fresh, or can spend some time on debugging, pnpm is amazing.
Why would one want to infer the requirements file? Would that be like `pip list --not-required --format-freeze`?
pip install -r requirements-dev.txt
Not the most convenient, but solves (works around) the issueHowever, like a sibling comment, I've also heard god things about Poetry. Probably worth giving it a spin somewhere and seeing if my initial thought still holds.
The USP of Hatch is that it uses the latest generation of Python standards and tech under the covers, so you have a unified tool that's less quirky than previous ones. Poetry and pipenv predate some of the improvements in Python packaging, so had to develop some things in their own way.
But it is very slow.
However, pdm and the likes are interesting because they are, indeed, PEP 582 compliants, which let you use __pypackages__ like one uses node_modules.
* https://iscinumpy.dev/post/bound-version-constraints/
* https://iscinumpy.dev/post/poetry-versions/
also pdm is developed by frostming, who's on PyPA (Python Packaging Authority) and more libraries are extending support for pdm (streamlit just today).
but in general i am not against poetry. it's a good tool and it pushed package managers a fair bit further - providing a superior alternative to conda imho