In Nix on the other hand, a particular hello package/binary will be dependent on the glibc that it is built with. That doesn't mean just the same glibc version, even if you'd change something like glibc's build flags it'd be considered a different version [1]. One of the benefits is reproducibility. If you build hello from the same nixpkgs revision as I do (barring any impurities), you'll have exactly the same hello with the same transitive dependencies.
Now, with hello this may not matter as much, because it's a simple program and glibc is fairly stable. But now imagine that you are doing some machine learning project and you want to provide a reproducible machine learning experiment. Building on a typical packaging ecosystem (eg. pip) is like building on quicksand. There are all kinds of subtle behaviors that can change and influence your experiment (eg. you machine learning library rewrites some compute kernels, and the small numerical differences pile up after a large number of layers). Though usually it ends at: my system has Python 3.11 and this experiment requires an old PyTorch version that was only built for Python 3.5 and earlier. The Nix approach gives you the benefit of pinning almost the whole world - your experiment and all its transitive dependencies.
Suppose you have a package repository like nixpkgs or Debian that contains tens of thousands of packages that are continuously updated. In Debian, if glibc is updated, it gets rebuilt and you are done. glibc has a stable ABI, so packages that were built against the previous glibc version don't need to be rebuilt. In nixpkgs if glibc is updated, every package that uses glibc needs to be rebuilt due to Nix' use of exact dependencies. And this is true for almost all non-leaf packages --- a package update requires all packages that are (transitively) dependent on that package to be rebuilt. And you need to store all these builds, so that users can download them, otherwise their machines need to rebuild the packages on every update. This is how you end up with a gigantic binary cache.
[1] This is a bit different now that Nix supports content-addressing for non-fixed output derivations.
Processed 974464 files, 327323 regular extents (751261 refs), 598287 inline.
Type Perc Disk Usage Uncompressed Referenced
TOTAL 40% 11G 28G 60G
none 100% 2.7G 2.7G 6.2G
zstd 34% 8.9G 25G 54G
Nix also supports automatic deduplication using hardlinks which I have enabled in this case.But i too wonder if they really have to keep everything around.
As mentioned, you doesn’t have to, but it has great value to be able to checkout any version.
and compression can do a lot for text files like sources etc. for example (last time i did that) the linux kernel source code was about 1.2GB uncompressed, but ~60-70MB in tar+gz format.
surely there must be a way to reduce that size.
I am inclined to argue that the fact NixOS is having this problem and discussion goes to show that you *still* have to be mindful of the resources you consume :)
While public cloud resources may seem infinite to some, at least the depth of the pockets of those shouldering the resulting costs should not.