Docker update ToS: Image retention limits imposed on free accounts
docker.com
docker.com
The out of the box registry does very little and has a very poor user experience.
But nowadays there are Harbor from VMware and Quay from RedHat that are open source and easily self-hostable.
We run our own Harbor instance at work and I can tell you... Docker images are NOT light. You think they are, they are not. It's easy to have images proliferate a lot and burn a lot of disk space. Under some conditions when layers are shared among too many images (can't recall the exact details here) deleting an image may result in also deleting a lot more images (and this is not the correct/expected/wanted behaviour) and that means that under some circumstances you have to retain a lot more images or layers than you think you should.
The thing is, I can only wonder how much bandwidth and disk space (oh and disk space must be replicated for fault tolerance) must cost running a public registry for everybody.
It hurts the open source ecosystem a bit, I understand... Maybe some middle ground will be reached, dunno.
Edit: I also run harbor at home, it's not hard to setup and operate, you should really check that out.
I can't see the bills, but we never worry about bandwidth usage and I am fairly sure that bandwidth is free, basically.
Keep in mind that since we run our own harbor instance, most of the image pulls happen within our openstack network, so that does/would not count against bandwidth usage (but image storage does). In terms of bandwidth thus, we can happily set "always" as imagePullPolicy in our kubernetes clusters.
Edit: openstack works remarkably well. The horizon web interface is slow as molasses but thanks to terraform we rarely have to use it.
FYI, Gitlab (free version) has a built in registry as well and it let's you define retention rules.
With harbor we can save docker images layers to switft, the openstack flavor of object storage (S3). That solves a lot of scalability problems.
AFAIK gitlab ships the docker Registry underneath so the problems stay, mostly. I think that harbor does the same. I skimmed the harbor source and it seems that it forwards http requests to the docker registry if you hit the registry API endpoints.
Haven't looked at Quay but as far as I know wherever there's the docker registry you'll have garbage collection problems.
One note on the side: I think that quay missed their chance to become the goto docker registry. Red Hat basically open sourced it after harbor had been incubated in the cncf (unsurprisingly, harbor development has skyrocketed after that event).
Scalability and garbage collection are actually two of the main areas of focus Quay has had since its inception. As you mentioned, most modern Docker registries such as Quay and Harbor will automatically redirect to blob storage for downloading of layers to help with scale; Quay actually goes one step further and (for blobs recently pulled) will skip the database entirely if the information has been cached. Further, being itself a horizontally scalable containerized app, Quay can easily be scaled out to handle thousands of requests per second (which is a very rough estimate of the scale of Quay.io)
On the garbage collection side, Quay has had fully asynchronous background collection of unreferenced image data since one of its early versions. That, plus the ability to label tags with future expiration, means you can (reasonably) control the growth problem around images. Going forward, there are plans to add additional capabilities around image retention to help mitigate further.
In reference to your note: We are always looking for contributors to Project Quay, and we are starting a bug bash with t-shirts as prizes for those who contribute! [1]
[1] https://github.com/quay/quay#quay-bug-bash
Edit: I saw the edit below and realized the bug bash is listed as ending tomorrow; we're extending it another month as we speak!
My message wasn't meant to dismiss quay, I hope it didn't come across like that.
I'll be giving a look at what's repo in during my vacation... I haven't loaded that many images on my private harbor registry, it might still make sense to switch :)
Edit: I just realized that the bug bash ends tomorrow... :/
I thought I'd just give some interesting background info and not so subtly ask people to contribute to the project :D
https://docs.gitlab.com/ee/administration/packages/container...
I think this is indirectly what people are complaining about. Having a free registry mitigates that. So they aren't far off track.
It's true we shouldn't be bitter about Docker. They did a lot to improve the development ecosystem. We should try to avoid picking technologies in the future that aren't scalable in both directions though.
For example, PostgreSQL works well for a 1GB VPS containing it and the web server for dozens of users, and it also works well for big sites. With MongoDB the VPS doesn't work so well.
Here's the conversation I saw:
https://stackoverflow.com/questions/33054369/how-to-change-t...
pointing to this:
https://github.com/moby/moby/issues/7203
and also there was this comment:
"It turns out this is actually possible, but not using the genuine Docker CE or EE version.
You can either use Red Hat's fork of docker with the '--add-registry' flag or you can build docker from source yourself with registry/config.go modified to use your own hard-coded default registry namespace/index."
"Fine we will just pay" - I have a personal account then 4 orgs, that's ~ 500 USD / year to keep older OSS online for users of openfaas/inlets/etc.
"We'll just ping the image very 6 mos" - you have to iterate and discover every image and tag in the accounts then pull them, retry if it fails. Oh and bandwidth isn't free.
"Doesn't affect me" - doesn't it? If you run a Kubernetes cluster, you'll do 100 pulls in no time from free / OSS components. The Hub will rate-limit you at 100 per 6 hours (resets every 24?). That means you need to add an image pull secret and a paid unmanned user to every Kubernetes cluster you run to prevent an outage.
"You should rebuild images every 6 mo anyway!" - have you ever worked with an enterprise company? They do not upgrade like we do.
"It's fair, about time they charged" - I agree with this, the costs must have been insane, but why is there no provision for OSS projects? We'll see images disappear because people can't afford to pay or to justify the costs.
A thread with community responses - https://twitter.com/alexellisuk/status/1293937111956099073?s...
This has been available as a docker image since the very beginning, which might not be good enough for everyone, but I think it will work for me and mine.
My guess is a lot of people don't know this. If it becomes better known, I can imagine something of an exodus from Docker Hub to GitHub Packages for OSS projects.
https://github.community/t/cannot-pull-docker-image-from-git...
It's crazy easy to do; just start the registry container with a mapped volume and you're done.
Securing, adding auth/auth and properly configuring your registry for exposure to the public internet, though... The configuration for that is very poorly documented IMO.
EDIT: Glancing through the docs, they do seem to have improved on making this more approachable relatively recently. https://docs.docker.com/registry/deploying/
Sometimes I don't push a new image version (if it's not critical to keep up with upstream security releases) for many months to a year or longer, but those images are still pulled frequently (certainly more than once every 6 months).
I didn't see any notes about rate limiting in that FAQ, did I miss something?
Honestly though 5 dollars a month isn't bad if you don't want to deal with hosting yourself.
> These limitations include but are not limited to [...] pull rate (defined as the number of requests per hour to download data from an account on Docker Hub) [...]
I read that as per account that owns the repository.
I'm not sure what protocol is used for pulling Docker images, but perhaps it could be enough to just initiate the connection, get Docker Hub to start sending data, and immediately terminate the connection. This should save bandwidth on both ends.
No, but they've got cash and are not price sensitive. Wringing money out of them helps keep it cheap and/or free for everyone else.
Enterprise customers might as well fork over cash to docker rather than shudder Oracle.
Can you patch a docker image? Sort of, but it's easier to rebuild. And that's what they do.
Hahahahahahaaaa!! No. Not in my experience.
Internally facing app that is AJAX glue over a legacy green screen app that is "only reachable from the internal network"? Probably not going get patched until something breaks.
Then your experience comes from somewhere with little concern for security.
Now, .NET Core 1.1 might not be the best example, but I'm sure you can think of some example.
But neither is storage
> have you ever worked with an enterprise company? They do not upgrade like we do.
I'm sure someone somewhere is going to shed a tear for the enterprise organisations with shady development practices using the free tier who may be slightly inconvenienced.
Set up CircleCI or similar to pull all your images once a month :)
https://www.docker.com/pricing/retentionfaq
What is an “inactive” image? An inactive image is a container image that has not been either pushed or pulled from the image repository in 6 or months.
It sucks that this isn’t for new images only. Now I have to go and retrospectively move my old images to a self hosted registry, update all my absolve scripts to the new uris, debug any changes, etc.
Good opportunity to sell a support contract. The point still stands - a six month old image is most likely stale.
Docker is doing the ecosystem a favor.
Neither is it for Docker...
A bunch of enterprises are going to get burned when say ubuntu:trusty-20150630 disappears.
It's not that they even have to rebuild their images... they might be pulling from one that will go stale.
I can't find this. It's not in the original link, is it?
If you have any ideas wrt to selecting important images, that'd be great.
A) "docker pull" commands and parsing the text that comes after it based on the command's syntax[1] to extract instructional references to images such as "docker pull ubuntu:latest, and
B) Searching for links/text beginning with "https://hub.docker.com/_/" to identify informational references to image base pages such as (https://hub.docker.com/_/ubuntu)
[1] https://docs.docker.com/engine/reference/commandline/pull/
The other issue is that (to my knowledge) the amount of papers on IA isn't terribly impressive. I think maybe indexing and going through SciHub will be better since some of these fields slap paywalls in front of their papers.
However, that's a pretty large task as well. The other thing is that papers rarely say "to reproduce my work do . . .". Usually the best we've got is a link to a GitHub repo (if that). I'm not sure how effective that strategy will be since it's guaranteed to be an under-count of the docker images we'd need to archive. Perhaps in conjunction with archiving all images that fall under particular search queries, we'd get the best of both worlds.
I've you've got ideas, feel free to hop onto efnet (#archiveteam and #archiveteam-bs) (also on hackint) to share your thoughts.
> You need an access token to publish, install, and delete packages in GitHub Packages.
https://docs.github.com/en/packages/using-github-packages-wi...
https://docs.github.com/en/packages/using-github-packages-wi...
GitHub Packages is different from GitHub Releases (and their artifacts) or cloning repos.
Yes, you do.
https://docker.pkg.github.com/v2/test/test/test/tags/list
> "code": "UNAUTHORIZED", "message": "GitHub Docker Registry needs login"
On the other side, I see the need for making money and storage/services cannot be free (someone pays somewhere for it - always), but 6 months is not that much for specific usages.
I wish they would grandfather images before this new ToS to not get wiped so that future images would be uploaded to more stable and accepting platforms so images on Docker Hub from research pre-ToS update don't get wiped.
START Run your own stuff on stuff you own. Run your stuff on other people's stuff you rent. This is too expensive to maintain at your rent. Pay us more. Back to run your own stuff on stuff you own. ...and so on, and so forth.
And this, ladies and gentlemen, is why anything worth doing is worth actually doing yourself. Nothing is worse than building something conditionally feasible on someone else only to have the rug pulled out from under you by sudden business pivots.
But that's the nature of the beast I suppose. I've certainly not found a great way to do it any other way.
For example:
docker pull python@sha256:2d29705d82489bf999b57f9773282eecbcd54873d308a7e0da3adb8a2a6976af
Pulls the latest Python image for Linux Alpine running on IBM System/z.
Not what you were asking, but if you have qemu-user-static installed, you can even run Docker images for other platforms. For example:
docker run -it python@sha256:2d29705d82489bf999b57f9773282eecbcd54873d308a7e0da3adb8a2a6976af
That's Python for Linux for IBM System/z. Check the CPU architecture:
import os;os.uname().machine
Running other platform's images under QEMU is going to be quite slow on a Raspberry Pi, I imagine (I haven't tried it). But of course for this case you don't have to run the image, you just want to pull it, so that doesn't matter.
I would not say rotting. From my perspective, the academic community has always lagged behind engineering best practices (except in their specific fields).
But that doesn't change the fact that it's just way easier to skim a paper and pull a docker image than follow every paper's custom build instructions and software stack.
It seems like you're arguing against using docker images, when docker builds solve the very issue you speak of.
Correct me if I'm wrong...?
Let's look at an example dockerfile for redis (based on [0])
FROM debian:buster-slim
RUN apt-get update; apt-get install -y --no-install-recommends gcc
RUN wget http://download.redis.io/releases/redis-6.0.6.tar.gz && tar xvf redis* && cd redis-6.0.6 && make install
(Note, modified from upstream for this example; won't actually build)The unreproducible bits are the following:
1. FROM debian:buster-slim -- unreproducible, the base image may change
2. apt-get update && apt-get install -- unreproducible, will give a different version of gcc and other apt packages at different times
Those two bits of unreprodicble-ness are so core to the image, that they result in every other step not being reproducible either.
As a result, when you 'docker build' that over time, it's very unlikely you'll get a bit-for-bit identical redis binary at the other end. Even a minor gcc version change will likely result in a different binary.
As a contrast to this, let's look at a reproducible build of redis using nix. In nixpkgs, it looks like so [1].
If I want a reproducible shell environment, I simply have to pin down its dependencies, which can be done by the following:
let
pkgs = import (builtins.fetchTarball {
url = "https://github.com/NixOS/nixpkgs/archive/48dfc9fa97d762bce28cc8372a2dd3805d14c633.tar.gz";
sha256 = "0mqq9hchd8mi1qpd23lwnwa88s67ac257k60hsv795446y7dlld2";
}) {};
in pkgs.mkShell {
buildInputs = [ pkgs.redis];
}
If I distribute that nix expression, and say "I ran it with nix version 2.3", that is sufficient for anyone to get a bit-for-bit identical redis binary. Even if the binary cache (which lets me not compile it) were to go away, that nixpkgs revision expresses the build instructions, including the exact version of gcc. Sure, if the binary cache were deleted, it would take multiple hours for everything to compile, but I'd still end up with a bit-for-bit identical copy of redis.This is true of the majority of nix packages. All commands are run in a sandbox with no access to most of the filesystem or network, encouraging reproducibility. Network access is mediated by special functions (like fetchTarball and fetchGit) which require including a sha256.
All network access going through those specially denoted means of network IO means it's very easy to back up all dependencies (i.e. the redis source code referenced in [1]), and the sha256 means it's easy to use mirrors without having to trust them to be unmodified.
It's possible to make an unreproducible nix package, but it requires going out of your way to do so, and rarely happens in practice. Conversely, it's possible to make a reproducible dockerfile, but it requires going out of your way to do so, and rarely happens in practice.
Oh, and for bonus points, you can build reproduible docker images using nix. This post has a good intro to how to play with that [2].
[0]: https://github.com/docker-library/redis/blob/bfd904a808cf68d...
[1]: https://github.com/NixOS/nixpkgs/blob/a7832c42da266857e98516...
[2]: https://christine.website/blog/i-was-wrong-about-nix-2020-02...
I was under the impression that Nix also wants to provide bit-for-bit reproducible builds, but that that is a much longer term goal. The immediate value proposition of Nix is ensuring that your source and your dependencies' source are the same.
In practice, most of the packages in the nixos base system seem to be reproducible, as tested here: https://r13y.com/
Naturally, that doesn't prove they are perfectly reproducible, merely that we don't observe unreproducibility.
Nix has tooling, like `nix-build --check`, the sandbox, etc which make it much easier to make things likely to be reproducible.
I'm actually fairly confident that the redis package is reproducible (having run `nix-build --check` on it, and seen it have identical outputs across machines), which is part of why I picked it as my example above.
However, I think my point stands. Dockerfiles make no real attempt to enforce reproducibility, and rarely are reproducible.
Nix packages push you in the right direction, and from practical observation, usually are reproducible.
What if it use a patched version of a weird library?
Software preservation is an huge topic and it is not done based on instructions.
Reproducible science is definitely a good goal, but reproducible doesn't mean maintainable. Really scientists should be getting in the habit of versioning their code and datasets. Of course a docker container is better than nothing, but I would much rather have a tagged repository and a pointer to an operating system where it compiles.
It's true that many scientists tend to build their results on an ill-defined dumpster fire of a software stack, but the fact that docker lets us preserve these workflows doesn't solve the underlying problem.
I haven’t filed any NSF stuff (yet) but didn’t come across any such hard requirements where you had to commit to something like zenodo or else to archive the result of your research work for archiving/citations purposes.
Also, if you do bio-type research, you can use Data Dryad too!
Two major issues I can see are old dependencies (pin your versions!) and out of support/no longer available binaries etc.
In which case, welcome to the world of long term support. It's a PITA.
https://docs.docker.com/engine/reference/commandline/image_s...
Many Docker based builds are not reproducible. Even something as simple as apt-get update failing with a zero exit code (it does this) adds complexity and most people don’t bother doing a deep dive.
Personally I use Sonatype Nexus and keep everything important in my own registry. I don’t trust any free offerings unless they’re self hosted.
You already have the source, now you have a dist.
It can take a Dockerfile and generate a 'locked' version, with dependencies frozen, so you at least get some reproducibility.
Disclaimer work for VMware; but on a different team.
I don't see any valid reason why anyone would upload and share a public docker image but not its Dockerfile and therefore do not pull anything from Dockerhub that doesn't also have the Dockerfile on the Dockerhub page.
A bonus to this is that you no longer have the risks of systems breaking because of Dockerhub or quay.io (which I haven't seen mentioned here yet, btw) being offline.
Having the images on dockerhub is more convenient, but as long as the paper says where to find the image this does not seem that bad.
Come on, if you stop just half a second and think about it, you know it is a stupid idea and you know that one day you will have a problem. You really don't have to be a genius for that. Same goes for all these other kinds of "services" that are bundled together with things that used to be a one-time purchase, like cars, etc.
Oh, I now have a t.v. that can play Netflix and Youtube, but is otherwise not extendible. But what happens in ten years? T.v. still works fine, but Netflix has gone bust and this new video-service won't work. Too bad, gonna buy a new t.v then. I can get really mad about this stupid short-sightedness everybody has these days.
Spoiler alert: one day Github will be gone too.
However, that is just information and that is not what I am talking about. I am talking about tools and things that stop functioning because they need some free service on the internet to work. Yes, all my projects and tools can work without internet access. Sure, they might not get updated anymore, but they will keep functioning and I could continue living my live no matter what shuts down.
This even extends to non-free services. For instance, I don't use Spotify, even though it is a nice product and I love exploring new music. But if there was a change of service, Trump decides to block my country economically or something like that, and I am kicked off the platform, I would suddenly have no music anymore, even though I would have paid for it for years. So I buy cd's and vinyl instead and rip them to flac.
I think it's a good idea to NOT be pulling from someone else's image on the internet.
It doesn't seem like a big deal really. It just means old public images from years ago that haven't been pulled or pushed to will get removed.
That means that "which image" is mixed with "where the image is". You can't deal with them separately. Because everyone uses the shorthand, Docker gets absolutely pummeled. Meanwhile in Java-land, private Maven repos are as ordinary as dirt and a well-smoothed path.
It's time for a v3 of the Registry API, to break this accidental nexus and allow purely content-addressed references.
`index.docker.io/foo/bar:latest` to be more exact, which is a URL, but not really a URI if we're being pedantic.
Docker doesn't really provide an interface to address images by URI (which would be more like the SHA), though in practice, tags other than latest should function closer to a URI
Some just pollutes my search result, I don't care that "yes, technically there's an image that does this thing I want, but it's Ubuntu 14.04 and 4 years old".
Even better, it prevents people from using these unmaintained image as a base for new project, which they will do, because many developer don't look at the Dockerfile and actually review the images they use in shipping product.
As a bonus perhaps this will mean that some a the many image of extremely low quality will go away.
I think it's fair, now you can either pay or maintain your images.
You also tend to have no idea what's in those images and what context people are creating them under. Sure, a lot of us know to check the dockerfile, github repo, etc but I have images with 10k+ downloads from OSS contributions but as you've said a whole lot of developers just grab whatever looks fitting on there. My biggest dockerhub pull has no dockerfile, no github repo, and is a core network configuration component I put up randomly just for my own testing because no docker image for it existed years ago.
I'm hopeful they'll add statistics/refers when this goes live.
Still, I keep my images mirrored on quay.io and I would recommend that to others (disclaimer: I work for Red Hat which acquired quay.io)
Non-abusive solutions include:
- extending docker to introduce reproducible image builds
- extending docker push and pull to allow discovery from different sources that use different protocols like IPFS, TahoeLAFS, or filesharing hosts
I'm sure you can come up with more solutions that don't abuse the goodwill of people.
It's already reproducible... sort of. All you need is to eliminate any outside variables that can affect the build process. This mainly takes the form of network access (eg. to run npm install, for instance).
Probably hold on long enough to get acquired?
Did they retain a significant amount of talent?
""" Can I use Quay for free? Yes! We offer unlimited storage and serving of public repositories. We strongly believe in the open source community and will do what we can to help! """
It wouldn't surprise me if people move to Github's registry for open source projects. https://github.com/features/packages
The egress pricing is going to be a dealbreaker. The free plan only includes 1GB out.
https://docs.gitlab.com/ee/user/packages/container_registry/
Actually just set up my own private registry and pull though registry.
Pretty easy stuff although no real GUI to browse as of yet.
This is all sitting on my NAS running in Rancher OS
It is cool seeing opensuse. Same the rancher
Keep free stuff free and add paid stuff. If your free stuff isn't sustainable, you really should have though that through early on.
This limit seems reasonable, because storage costs are expensive. But it should have been implemented day one so people have reasonable expectations on retention. Other's have mentioned open source projects and artifacts for scientific publication being two niche use cases where people still might want this data years later, but it'd be rare for it to be pulled every six months.
I only have a few things on docker hub, but I'll probably move them to a self-hosted repo pretty soon. At least if it's self hosted, I know it will stay up until I die and my credit cards stop working.
Its not an unreasonable strategy to provide generous free hosting if you derive some other business benefit from it (YouTube being another example).
But Docker Inc. found their moat was not that deep and other projects from the big cloud providers killed the market they saw for Docker Enterprise and they sold it off.
So now they just have docker.com and Docker CE - which even that has alternatives now with other runtimes existing. So they need to make docker.com a profitable business on its own or find something else to do which changes the equation significantly.
* Images are just files. You can copy them around (or archive them) like any other file. Docker's layer system is cool but brings a lot of complexity with it.
* You can build them from Docker images (it'll even pull them directly from Dockerhub).
* Containers are immutable by default.
* No daemon. The runtime is just an executable.
* No elevated permissions needed for running.
* Easy to pipe stdin/stout through a container like any other executable.
Never heard of Singularity before, and it does look interesting. Wanted to point out though that you can create tarballs of Docker images, copy them around, and load them into a Docker instance. This is really common for air-gapped deployments.
[0] https://www.pluralsight.com/courses/implementing-self-hosted...
> An inactive image is a container image that has not been either pushed or pulled from the image repository in 6 or months.
>
> How can I view the status of my images
> All images in your Docker Hub repository have a “Last pushed” date and can easily be accessed in the Repositories view when logged into your account. A new dashboard will also be available in Docker Hub that offers the ability to view the status of all of your container images.
That still does not tell the whole story, does it? I still don't know if my image have been pulled for the last six months. Only when I pushed it.
Best to assume the worst but still plenty of time to write a cron that pulls all your images, assuming for some reason you need images you don't pull for > 6 months.
Registry is open source (https://github.com/docker/distribution) and implements the OCI Distribution Specification (https://github.com/opencontainers/distribution-spec/blob/mas...) if you want to dig into it.
In theory you could modify the spec/application to try to break layers down into smaller pieces but I have a feeling you would reach the point of diminishing returns for normal use cases pretty quickly.
> Containers are increasingly used in a broad spectrum of applications from cloud services to storage to supporting emerging edge computing paradigm. This has led to an explosive proliferation of container images. The associated storage performance and capacity requirements place high pressure on the infrastructure of registries, which store and serve images. Exploiting the high file redundancy in real-world images is a promising approach to drastically reduce the severe storage requirements of the growing registries. However, existing deduplication techniques largely degrade the performance of registry because of layer restore overhead. In this paper, we propose DupHunter, a new Docker registry architecture, which not only natively deduplicates layer for space savings but also reduces layer restore overhead. DupHunter supports several configurable deduplication modes , which provide different levels of storage efficiency, durability, and performance, to support a range of uses. To mitigate the negative impact of deduplication on the image download times, DupHunter introduces a two-tier storage hierarchy with a novel layer prefetch/preconstruct cache algorithm based on user access patterns. Under real workloads, in the highest data reduction mode, DupHunter reduces storage space by up to 6.9x compared to the current implementations. In the highest performance mode, DupHunter can reduce the GET layer latency up to 2.8x compared to the state-of-the-art.
You now have that, with docker.
0: https://aws.amazon.com/free/?all-free-tier.sort-by=item.addi...
that would still pose a problem, not cost-wise, but you'd still need to download image after image. Will a single instance be capable of "downloading" all existing images every 6 months?