Linux package managers are slow
michael.stapelberg.ch
michael.stapelberg.ch
I think I haven't been able, in the last years, to update a Windows 7 SP1, without using hacks.
AFAIR (not 100% sure) some Windows Update fixes have been released after Windows 7 SP1 (which is the last available image), which means, in order to install a fresh system and update it, one has to install the O/S, then patch, then update.
So this explain at least some cases. Although, I remember that even after patching, WU would hang.
Anybody ever used the peer to peer local caching ?
Yes, I never understood why this was never improved across the years (and systems).
The only explanation with which I could come up, is that the team in charge is still the same team who designed progress bars in DOS and Win 3.1 times, and they have been granted the role of tradition keepers.
And at the end of this long and ardouous process, you only have an updated OS and still have to manually (or use a tool like Chocolatey) update your non-MS software.
It baffles me that this as considered acceptable considering these issues are still present even in domain environments with local WSUS servers.
Or NixOS and Guix that allow you incredibly fine-grained control through configuration? (Pinning multiple versions down to particular hashes.)
I don't think that's true at all. I choose when my kernel gets updated, and which dkms modules I want to update (as some of them are a bit... bleeding edge). I choose which of my apps needs an update.
I can choose just security, or even 'just do this particular security update for this particular program'.
This ability does not exist in Linux package managers since they deal with packages - there is only one newer package to install, and installing that new package replaces the entire existing package. Hotfixes work by applying patches to the affected components' files, which is why they can be independent of other hotfixes applied to the same components.
Hotfixes are also grouped into larger updates, the largest of which is a service pack. (You released N hotfixes over the last month and expect everyone to apply them in the future. There's no point making every user in the future download N updates where one batched up one would suffice.) So when Windows Update has to answer the question "Is this new update required or is it already installed via a previous update?" it has to do a bit of gymnastics to figure out the answer.
So both approaches have their pros and cons. The hotfix method means you can get targeted fixes with no feature updates and smaller chance of regressions, at the cost of being more complex to solve than "if server's version > installed version, download and overwrite".
Since it only affects Foo users that are also BFW users, it might happen that the hotfix happens to break Foo users who have QuuxCorp Internet Optimizer installed instead. The testing might not have caught it since the primary purpose of the testing was to ensure that the issue is fixed for BFW users, with minimal sanity testing that it doesn't break non-BFW users.
So QIO users can mitigate themselves by uninstalling the hotfix. They didn't need it anyway. Eventually another hotfix may be released that supersedes the first one and is tested for both BFW and QIO, and this latter hotfix will eventually be bundled up into a general Foo update.
All this happens independently of another Foo hotfix released after the BFW one, to fix the issue where it crashes when you open a .txt that's actually a .bmp.
The equivalent with Linux packages would be to have .1, .2, .3 package releases, but QIO users can't can't keep the changes in .3 and remove the changes in .2.
The equivilent of Windows SxS would be like running everything in it's own docker container. Assuming each one was updated concurrently & they used a shared package cache this could still be quite fast compared to Windows.
Even accounting for Window's 'impressive' ability to let you install any combination of updates/hotfixes. It's basic software engineering that the common case should be fast and considering that in 99.99% of Windows instances are going to have all of the updates installed, it's very apparent that the Windows update system & SxS were pooly designed and is an major area at Microsoft suffering from brain drain. I feel the most likely case is that no one currently at Microsoft (and still working as a programmer as many of the old talent are either retired or upper executives) understands it enough to fix the issue.
WinSxS is not about hotfixes either, but for having different versions of components simultaneously installed.
Neither of these are relevant to the point of this subthread.
Look for `max_parallel_downloads` here:
As for volunteer servers, if a particular server can't handle >N connections, it can drop them. No reason to gimp clients instead.
And specifically in openSUSE's case, nobody has put forth server bandwidth / limits as a reason not to parallelize downloads. [1][2][3] It's just not been done because nobody (including myself) has cared enough to figure out how to do it.
[1]: https://features.opensuse.org/307862
- The QEMU results are highly misleading, because there are tons of different configurations possible with QEMU. Nix builds all QEMU architectures while the Alpine one only is getting qemu-system-x86_64. It's apples to oranges.
- The Nix docker image is not actually running NixOS, but instead running Alpine. It just has the Nix package manager installed on top of that.
When I was on Ubuntu though, which does have version requirements taken into account in packages, I didn't find apt to be slow either.
You know what's slow? The conda package manager. Or installing software - practically any software - on Windows. I regularly budget half an hour to an hour just to install something on the Windows computers I'm forced to use.
Now brew and conda, they often take minutes to solve for the environment, which is what I think of when you say "slow". That part is probably mostly algorithmic, and it should be fast, unlike network and disk.
$ time xzcat /var/cache/pacman/pkg/cuda-10.1.243-1-x86_64.pkg.tar.xz >/dev/null
1:17,33 total
Just decompressing that 1.8GB package takes 77 seconds.There was a discussion in April to switch compression to zstd - which reduces the decompression time for that package to only 4 seconds (!) while keeping the same file size. That way the disk is the bottleneck again, as it should be. That switch would be awesome, but it looks like it's not happening currently.
I would also be interested to see if it’s using parallel decompresssion when possible. I have an intuition it doesn’t, since I have always found Arch defaulting to single thread decompression, e.g in makepkg or for compressing with mkinitcpio.
2. Even if you don't, why would you have to maintain a separate container? If `RUN apt install jq` is indeed the second entry in your Dockerfile after the `FROM ...`, then it should be cached and reused indefinitely -- that's half the point of Docker in the first place. It won't work if you have a command like `COPY` before the `RUN apt`, though.
Yesterday I had a nearby arch mirror fail and timeout. And then the secondary backup mirror downloaded things at 20M/s.
Also. just tried qemu on arch and there definitely wasn't 124mb of data transferred. (Total download size: 7.85mb) so this can vary heavily depending on which deps you have pre-installed for something elsewhere.
Doesn't seem to be the best methodology.
nix-env -iA nixpkgs.ack
nix-env -iA nixpkgs.qemu
This is because Nix considers an entire tree of Nix expressions at once, rather than examining each package in isolation. When nix-env is run without the -A flag, it must evaluate the entirety of nixpkgs and then search for the desired package amongst the complete tree. With -A instead, it is possible to evaluate only the portions of the tree which are relevant to the named package. $ time nix-channel --update
unpacking channels...
created 2 symlinks in user environment
nix-channel --update 1.92s user 16.58s system 15% cpu 1:56.62 total
$ time nix-env -iA nixpkgs.ack
installing 'perl5.28.2-ack-3.0.2'
these paths will be fetched (0.06 MiB download, 0.19 MiB unpacked):
/nix/store/43rvd2kinid0g48ng07d782dsqb497g1-perl5.28.2-ack-3.0.2
/nix/store/k3f819h2ncadkd6g8yzyzg24wn3vr609-perl5.28.2-ack-3.0.2-man
/nix/store/p82srjihmh7vnq4szl9q83adcl9r6j2v-perl5.28.2-File-Next-1.16
copying path '/nix/store/k3f819h2ncadkd6g8yzyzg24wn3vr609-perl5.28.2-ack-3.0.2-man' from 'https://cache.nixos.org'...
copying path '/nix/store/p82srjihmh7vnq4szl9q83adcl9r6j2v-perl5.28.2-File-Next-1.16' from 'https://cache.nixos.org'...
copying path '/nix/store/43rvd2kinid0g48ng07d782dsqb497g1-perl5.28.2-ack-3.0.2' from 'https://cache.nixos.org'...
building '/nix/store/p765c3dxn32zccl4q9xjgr74n2ljgiky-user-environment.drv'...
created 134 symlinks in user environment
nix-env -iA nixpkgs.ack 0.19s user 0.24s system 36% cpu 1.158 totalOn my machine, for the original command I get:
created 56 symlinks in user environment
real 0m 25.64s
For your command I get: created 56 symlinks in user environment
real 0m 4.39s> Pain point: too much metadata
and
> I expect any modern Linux distribution to only transfer absolutely required data to complete my task.
Consider the case of searching for a package. That search can be performed locally and you would need the metadata locally or it can be performed remotely and the results sent over the wire. It has to be performed somewhere.
If performed remotely than the remote endpoint knows what you've been searching for. If you are trying to perform a search across multiple package repos how do you to that securely? How do you not leak information? The remote endpoint can capture, log, and use information about who searched for what. Then tie that back to company and other data. There's a lot of sensitive corporate information that can be contained in these searches. So, do you do perform searches remotely or locally? Many systems choose locally but that entails downloads a dataset.
There are a bunch of things like this to take into account.
Other things, like serial downloads, are likely a product of the package managers age. I would hope modern ones work out parallel downloads.
I never looked into the technical details of yum/dnf metadata, but I suspect there must be a much smarter way to manage that stuff. Incremental updates for instance, or maybe just a different format (is it not xml?).
If you use nix then you'll know that a “nix-shell -p” is basically the url-bar of the browser.
If a browser is a VM in your OS then why not run distri in a VM and speedily launch native apps, horribly useless example.. for now.
If the browser is eating the world anyway then incentives to have native apps be incompatible between operating systems fade away and throwaway envs like nix-shell (and to a lesser extent nixos-shell) become even more interesting. . So long as they are fast enough to feel as effortless as opening a web page.
Both fetch and install are useful to know, but for very different reasons. Only so much of fetch speed is in control of the distro setting it up (but you can set up your own mirror if you like...), but install speed once the files are already staged locally is also useful to know. Additionally, whether the package manager keeps a local cache of metadata for packages, and whether it checks and updated it every request or only at certain intervals or when requested will affect this as well (and is very important, if you want to know about security updates immediately and not after some interval of hours).
What the author has done is the equivalent to testing how fast certain one mile and 10 mile stretches of highway are, but without controlling for time of day, and not realizing rush hour is invalidating all the results. The first step to measuring something is to learn about what you're measuring so you can avoid issues like these.
Also, without separating out the time spent waiting for the download, the author's choice of mirrors probably has a significant impact.
0: https://www.archlinux.org/packages/extra/x86_64/qemu/
1: https://pkgs.alpinelinux.org/package/edge/main/x86_64/qemu
If Linux package managers were faster then <it would affect me how?>
You can customise the sources, pay attention to what you want installed/upgraded... or even run a mirror
Someone else mentioned concurrent downloads which sounds like a good idea, otherwise... thankful.
dnf install /usr/share/mypackage/file
However they are not downloaded if you're just installing a package by name (even for dependencies) which is what this test is all about.Otherwise 20-60 seconds saved doesn't matter to me personally for a one-time event. I'm guessing it doesn't matter to the package-manager developers either.
Let's say there are 30 million Linux users who are wasting an hour every year on slow package managers. 30 million hours means 3424 years wasted every year or the total average lifespan of 48 people. Package managers kill 48 people per year. ( /s )
(There's some related Steve Jobs quote here somewhere.)
It's odd that Alpine lets you install e.g. "emacs-nox" on an IoT device connected via crappy LTE, on the other side of the world, faster than you can install it on your local Ubuntu dev machine, with like 100-1000x more network bandwidth, 10-100x more CPU and 100-1000x faster local storage.
It'd be awesome if updating a node was a sub-second operation, it's definitely a big benefit to cloud based computing.
Alpine is fastest not just because it has a more efficient metadata format, but also because it has far fewer packages than the others.
At least the ones where the generic container is Ubuntu/Debian based. `apt update && apt install -y` is really slow. But alpine based containers with the equivalent `apk install` are blazing fast.
Consider the amount of time it takes to update a system vs. its overall runtime. Now consider what's important in a package manager. Consistency, rollback, configuration management, audit trails, etc. The speed of the package manager is at the absolute bottom of things to be concerned about.
zx2c4@thinkpad ~ $ genlop -t qemu
Sat May 18 15:36:44 2019 >>> app-emulation/qemu-4.0.0-r2
merge time: 6 minutes and 49 seconds.