Just imagine the vast number of poorly cached CI jobs pulling gigabytes from Docker hub on every commit, coupled with naive aproaches to CI/CD when doing microservices, prod/dev/test deployments, etc.
Just imagine the vast number of poorly cached CI jobs pulling gigabytes from Docker hub on every commit, coupled with naive aproaches to CI/CD when doing microservices, prod/dev/test deployments, etc.
As far as financially, Docker now charges for enterprise use of Docker Desktop, which we've also started paying for. But I'm sure the bandwidth for running Docker Hub isn't cheap.
Which breaks all normal caching proxies that can be easily used.
The popular packaging formats are better at this: distribute keys over https, then just http to download and verify packages. Squid and co work just fine here.
Once quotas come in, you’ll find all the tooling and guides make it super simple.
I don't know about that. One of the benefits of introducing something like Nexus can be a noticeable speedup for your own builds, once your proxy repositories have the versions of dependencies you need cached.
Of course, you could use some sort of a local build cache (e.g. m2 directory for Maven packages on the server) but when you're building your own containers it doesn't always turn out to be as viable, especially when you have many CI nodes, each of which would have a local build cache.
On an unrelated note, something like Nexus also gives you the ability to easily start deploying/using your own container images, libraries or even arbitrary files, all without having to store your data in the cloud, or figure out how many different solutions/accounts you might need (otherwise you'd need some packages on Docker Hub, some Node packages on npm, some Java packages in Maven repos etc.).
As far as I know there aren't any projects doing peer-to-peer distribution of container images to servers, probably because it's useful to be able to use a stock docker daemon on your server. The Kraken page references Dragonfly [2] but I haven't grokked it yet, it might be that.
It has seemed strange to me that the docker daemon I run on a host is not also a registry, if I want to enable that feature.
It's also possible that in practice you'd want your CI nodes optimized for compute because they're doing a lot of work, your registry hosts for bandwidth, and your servers again for compute, and having one daemon to rule them all seems elegant but is actually overgeneralized, and specialization is better.
Steam never used P2P downloading but they've had deals with ISPs and communities really early and made their own "CDN" full of caching servers everywhere, so if you waited a few hours to update, it might make it to your local server in time. Now I think they're on Akamai as well.
Docker requires to pull a 1GB image to run a 1MB binary.
Increases of 100x to 1000x are pretty common and moving data around is quite harmful.
gcr.io/distroless/static-debian11 latest 10d0cc57bead 2.36MB
No functionality beyond running a binary, but, hey, that's what GP wanted.I hit the rate limits that others talk of in the comments, which motivated me to use Nexus for both proxying and storing my own container images. I actually tried using GitLab Registry before, but something with a bit more flexibility (cleanup policies) is nice! So far, it's been pretty good, I actually wrote about the process on my blog, "Moving from GitLab Registry to Sonatype Nexus": https://blog.kronis.dev/tutorials/moving-from-gitlab-registr...
Another thing that I tried, however, was to only rely upon Docker Hub for the base images that I want (Ubuntu in my case) and then build everything I need on top of that, doing things like installing Java/Node/Python/Ruby/... manually, adding utilities I want across all of the images etc. Once again, I wrote about it on my blog, "Using Ubuntu as the base for all of my containers": https://blog.kronis.dev/articles/using-ubuntu-as-the-base-fo...
That approach is absolutely more work, but also is something that's underexplored and works really nicely for me. Now I mostly rely on the OS package manager repositories (or mirrors of those), put less load on Docker Hub, don't risk running into its rate limits and also have common base layers across most of the images that I build, which in practice means less data actually needing to be downloaded to any of the servers where I want to utilize my images.
Of course, the downside is that getting something like PHP running was an absolute pain (tried with Apache, didn't work for some reason, then moved over to Nginx), and I technically miss out on some of the more complex space optimizations because if you look at the Dockerfiles for some of the more popular images, like OpenJDK, you'll occasionally see some interesting approaches, like getting the software package as a bunch of files and "installing" them directly, as opposed to using something like apt/yum: https://github.com/adoptium/containers/blob/08dd7d416cee0fe0...
Then again, personally I'd much prefer to rely on packages that I can get from something like apt directly, even if some of those versions can be a bit older (or add the project's official apt repositories as needed). One exception for software I don't install myself: databases, due to how good the majority of images out there (e.g. MySQL/PostgreSQL) are already. I just proxy/re-host those with different tags as needed.
> We've since started caching our images locally using Sonatype Nexus Repository Manager plus hosting our own registry for some simple things we used to be pulling from Docker Hub.