How to own your own Docker Registry address
httptoolkit.com
httptoolkit.com
The K8s proxy is redirecting from only hosting on GCR to community-owned registries - https://kubernetes.io/blog/2023/03/10/image-registry-redirec...
You can view the code here - https://github.com/kubernetes/registry.k8s.io
But because everyone is already pointing at gcr.io (just like many openfaas users point at docker.io/) - they're having to do a huge campaign to announce the new URL - the same would apply with the author's solution here.
I wrote some automation for hosting (not redirects) in arkade with the OSS registry - Get a TLS-enabled Docker registry in 5 minutes - https://blog.alexellis.io/get-a-tls-enabled-docker-registry-...
The registry is also something you can run on a VM if you so wish, and have act as a pull through cache.
Apart from reliability - GitHub's container registry is the current next best option - but we have to ask ourselves, what happens when they start charging or the outages start to last longer or are more frequent than 1-2 times per week as we've seen in Q1 2023.
Did people forget about Google Code repositories already?
I cannot for the life of me understand why people use google products like google isn't going to shut them down whenever they want to. Stadia was the most amazing example.
"Surely, google wouldn't do that. Stadia customers are paying customers."
3 years from launch to shut down announcement. 3 months from shutdown announcement to actually shutting down and losing all your investment in Stadia.
I don't understand how long google can keep getting away with this.
A slightly fancier version of this concept is the Kubernetes official registry, registry.k8s.io .
https://github.com/kubernetes/registry.k8s.io
Their registry forwards to a container registry by default. But if it detects if the request is coming from inside AWS or Google Cloud, it forwards requests for blobs to a S3 or GCS bucket near the requester. This saves money on cloud egress charges.
I would've hoped that at this point we would have a true decentral solution for this sort of thing. Despite all the blockchain/dapp/web3 hype for many years they have no practical solution for anything.
We have all the pieces it seems, torrent and dht/magnet links work, ipfs works, web of trust works. And yet we don't seem to manage to work collectively on true decentral solutions to the issue of centralization of critical internet infrastructure. Why can't we all work collectively together and share resources so we aren't dependent on the whims of some shaky businesses, we are all constantly at risk of them turning on us for profit.
All the cutesy technowords-of-the-day, blockchain/dapp/web3/torrent/magnet links, are just a bandage over a greater point: ever since the atomic bomb, once we became able to destroy the planet, we needed to become a new species, evolve our cone of care. We were unable to do so and hopefully we will be extinct before we destroy the planet, let some other species have their try in a few million years, before the sun runs out.
I would also draw a distinction between web3 and torrent technologies. Torrents work great, and it doesn't even give its users a monetary incentive to seed, people do it anyway. But web3 makes everything transactional and builds everything around individualist monetary incentives, and yet no useful application was ever (so far) conceived by it. So perhaps torrents and the wikipedia (and similar projects) work because it doesn't built everything around the free market libertarian fever dream.
Perhaps I am too doomy, but as we see every day, and now with the GPT advances, almost every hour, a bridge being built between the information space, the decision-making space, and the 3D space of the physical world, and this bridge being restricted to only certain entities, it makes one wonder: would a Wikipedia even be possible today?
[1] Diamonds are a De Beers invention and a monopolistic violent endeavour, moissanites are cheaper, no artificial scarcity, and better looking https://en.wikipedia.org/wiki/Moissanite
Wikipedia works thanks donations.
Voluntary donations are the free market libertarian equivalent of involuntary taxation.
Web3 doesn't work because there's not enough value being provided, party because paying micro transactions is too unfriendly. That's a hard problem and it's lack of success is a clear demonstration of the market working as intended and not rewarding something useless.
The socialist equivalent is a government owned web3 which we all have to pay with taxes and we're increasingly close to getting this.
Eh, we can take solace in the fact we still don’t have the capability to destroy the planet. Vastly alter the current environment, cause mass extinctions, and irradiate the planets surface, sure we can do those things. But the planet won’t care, and life, well, life finds a way.
Sadly that ship has sailed, until a dollar/government backed blockchain which allows to do such things will pop up. Which won't happen i think
Image registry over ipfs would solve the problem of decoupling who actually stores the image from the process of discovering the locations
https://github.com/containerd/nerdctl/blob/main/docs/ipfs.md
But yeah, any system where you store blobs of other people that you yourself don't want is potential liability
"Just" docker registry proxy that had torrent support would be fine enough solution for the distribution. But good luck convincing anyone in ivory tower of security that opening some random ports to entire of the internet is a good idea
Tragedy of the commons at its finest.
This is not correct. It's the "organization" features which are going away. That is the feature which lets you create teams, add other users to those teams, and grant teams access to push images and access private repositories. Multiple maintainers can still collaborate on publishing new images through use of access tokens which grant access to publish those images. It's kind of a hack, but it works. You would typically use these access tokens with automated CI tools anyway. This will require converting the organization account to a personal user (non-org) account. (Interesting note/disclosure: I was the engineer who first implemented the feature of converting a personal user account into an organization account some time around 2014/2015, but I no longer work there.)
For open source projects which are not part of the Docker Official Images (the "library" images [1]), they announced that such projects can apply to the Docker-Sponsored Open Source Program [2].
I would also heed the warning from the author of this article:
> Self-hosting a registry is not free, and it's more work than it sounds: it's a proper piece of infrastructure, and comes with all the obligations that implies, from monitoring to promptly applying security updates to load & disk-space management. Nobody (let alone tiny projects like these) wants this job.
Having most container images hosted by a handful of centralized registries has its problems, as noted, but so does an alternative scenario where multiple projects which decided to go self-hosted eventually lack the resources to continue doing so for their legacy users. Though, I suppose the nice thing about container images is that you can always pull and push them somewhere else to keep around indefinitely.
[1] https://hub.docker.com/u/library [2] https://www.docker.com/community/open-source/application/
I don't think this is possible.
Sure you could implement a finer-grained deduplication or transfer mechanism, but I doubt this would scale as well. Many large image layers consist of lots and lots of small files. The overhead would be tremendous.
How? Well, the most simple way is compute the digest of the content and look it up, oh wait :thinking:
Perhaps there is some glimmer of hope for self-hosting after all.
The last time I've looked at self-hosted CI/CD, Concourse stood out as one of the more promising options.
As for code and container registry - Gitea? It seems like it has an integrated container registry now, so that's a plus.
GitLab is an overbloated mess that you can't really justify unless you have organization-style funding/tax-writeoffs for the server (at which point it's easily the best choice). It expects CI/CD to exist for all projects, even though it can run without it, you'll be missing quite a number of features (the main example is that GitLab demands release builds to be generated through CI unless you want to manipulate the API with curl on your dev machine, vis-a-vis uploading things). It needs a somewhat beefy server unless you go out of your way to downtune the entire thing (which requires quite a bit of configuration), a 5$ VPS will not suffice.
Gitea mostly rocks and in my experience runs on even that 5$ VPS, but it does not ship with any CI/CD by design. They do have a list of external services[1] that can provide CI that can integrate with their software, so you can have CI/CD. Personally I'd recommend Gitea if you're looking to selfhost.
[0]: https://docs.docker.com/registry/
[1]: https://gitea.com/gitea/awesome-gitea#user-content-devops
Currently using Gitea + Sonatype Nexus + Drone CI which has worked nicely for my own needs, after previously running self-hosted GitLab for a few years, but finding the updates to be a bit problematic: https://blog.kronis.dev/articles/goodbye-gitlab-hello-gitea-...
That said, Woodpecker might be a more open CI offering, licensing wise and works similarly to Drone.
But personally I wouldn't judge others for picking whatever else they are familiar with and what works for them, even if that choice would be Jenkins or something like that.
Self-hosting your own private registry is easy and cheap but for a public registry you’re essentially writing a blank check letting randos across the internet pull your image a million times from their CI pipeline and cost you egress fees.
Probably not surprising given that this is the blog of HTTP Toolkit, but instead of debugging the HTTP requests, they could have gotten much of the same information by reading the introduction of the API docs (https://docs.docker.com/registry/spec/api/#overview).
- there are such docs
- you can find them
- they won't lie to you
- they'll be enough to give you a picture of how everything works
It's great that Docker fits this criteria, but in general tools that let you cut out a few of those steps and inspect how the actual thing is running live are also nice.Ideally, consider using both: trust the docs, but validate regardless to have more confidence.
HashiCorp uses a service discovery protocol for Terraform: https://developer.hashicorp.com/terraform/internals/remote-s...
It allows domain owners to use a "pretty" name for their Terraform services (e.g. terraform.ycombinator.com) while pointing to a different host and path (e.g. an S3 bucket).
Decoupling the "root" of a registry from where its API is implemented is a great layer of indirection I wish existed in the container image registry ecosystem without speciaised server software to perform API-aware redirects.
I'm not sure about them, but Nexus might fit the bill from that list: https://github.com/caprover/one-click-apps/blob/master/publi...
It's what I'm using for myself (though with just Docker Swarm + Apache2, without Caprover) and has worked well for years.
https://caprover.com/docs/app-scaling-and-cluster.html#setup...
I'm using official registry, docker-registry-ui for auth and web-based repository browser (which essentially proxies the docker API requests to registry, while providing web access to list the images and nginx-proxy-manager in front of all this (needed as the NAS I'm running this on hosts some other stuff as well).
Fact: Docker wants to recover the cost of hosting all those images. There is a cost to storing each image. (is that right?)
Fact: Docker can easily change their API and how they handle redirects to ensure this scheme does not work now and does not work in the future.
If the goal is survive the next Dockerpocalypse this seems unlikely to do it. Or perhaps I misunderstand.
Is Docker doing such a poor job that this is still misunderstood?
Wholly non-commercial open source projects. Meaning, for example, even if you sell consulting services for your project on the side, you don't qualify. And it's re-evaluated every year.
The solution in the article allows such projects to use their own domain without incurring any of the maintenance burden of running a public registry.
One of the problems for open-source projects currently is that any move (even to another provider) causes a lot of disruption, because all users of the project need to update the address they pull from. This solves that.
That feels a bit snarky. The article opens with that path, but concludes it's heavier than they want. Then suggests an option that's lighter/thinner and meets their needs.