Docker to rate limit image pulls
docker.com
docker.com
In most cases it would seem $5/user isn't much as a business or organizational expense, and if it's a personal project 200 images in six hours seems pretty solid?
I'm just sort of shocked it's so cheap. I figured it'd be like $25-$100 a month or something just because of all the bandwidth someone could probably burn building random/broken shit over and over.
I'm mostly curious because I've considered using Docker recently for personal projects and my home server; but I'd rather not invest a bunch of time porting things only to find out it's gonna be way more than $5 (not including my time).
Or is it that hard to just build your own images? To be honest, I've never really understood that part of Docker (using other folks images)... It always seemed like an enormous security risk[1]. FWIW, I still deploy my personal projects using chef/ansible, shell scripts, and systemd units like some sort of curmudgeon-y monster...
[1] Everything is a security risk (APT, source, etc), I know, I get it. No one need scribe miles of pedantry into the comments explaining it to me. It's what each of us has tolerance for and can accept that matters.
"Stock CentOS was good enough before and it's good enough now!"
It's hard enough to save money as it is.
News websites blow my mind with this - if I forked over $5 to every news outlet I occasionally like to read, I'd be spending at least $500 maybe more per year JUST to get access to some random person's biased recant of what's happening in the world. If there were a news source that did the opposite of this, and basically provided a bullet list of objective, non-biased events boiled down to exactly what I need to know, that might be something I'd pay for. Hell, it would save you time over filtering the opinionated BS out.
The perpetual spam is worse for me than the $5.
If you get your news like most people, via link aggregators like HN, Reddit Facebook, Twitter then you get linked to dozens of publications that all want a $5/mo. commitment which is untenable.
But... in their news app, there's always a couple of interesting articles I might want to read, but I'm not signing up. They have my info, and a Touch ID device I'm holding tied to my payment info. "Read this article for 50c?" I'd certainly give some a read now and then.
I have been subscribing valuable information sources since 1995, so I do get the point.
Also there are news wire services that do mostly what you're describing. If you just want to be entertained (most people read news for entertainment), then they don't really care about the facts. They want to hear about so and so blasting so and so or whatever. But if you're trying to make money from information (traders, journalists, etc), then you really don't want to be reading the kind of stuff the New York Times is publishing.
Just a subscription to The Information is $399/yr. Add a subscription to the Times, and you've already blown past your $500 budget.
Consider for example the current riots going on in the US. How do you objectively report on that? With bias, on one side you have "peaceful protest disrupted and escalated by the police", on the other side you have "police intervening in riots to maintain order and protect property". There's not really inbetween.
Done.
Much as we'd all love an "unbiased" news source, the reality is that bias is a very hard problem to solve well.
If one side is lying, objective reporting would tell you which side it was.
If you’re deploying to a cluster with 200 machines, you could easily hit this if you use the public registry though. However, if you’re managing that size cluster you can probably afford the fee, but more importantly, you should probably pull once to a local registry and use that to deploy to your cluster anyway.
If you do something like this, you absolutely MUST have a local registry.
Harbor [1], JFrog [2], and Quay [3] would be the first ones that I look at.
Harbor is open source, free, and a member of the CNCF. You will need to do a little bit of work to set it up to scale properly. JFrog offers a SaaS registry, but you will pay big $$ based on pull traffic. Their commercial site license is about $3k/year. Quay is older than either of them, stable, and high quality. I'd start with Harbor these days.
[1] https://goharbor.io/ [2] https://www.jfrog.com/confluence/display/JFROG/JFrog+Artifac... [3] https://quay.io/
I have pulled and run 20k times a 1GB image in less than 10-15 minutes without breaking a sweat.
Finally GitHub packages offers a registry out of the box . It is great for CI and devs to access . I generally have the tags mirrored from tags GitHub for production to ACR .
That said, word of warning for anyone looking at GitHub Packages for docker registry: it's broken with containerd and some other similar tools. They (GitHub) are currently working on a fix: https://github.com/containerd/containerd/issues/3291
1) It is broken and unusable on Kubernetes and Docker Swarm.
2) It is flaky often returning 500 type errors.
3) It is expensive as the amount of pull bandwidth is very limited.
Hmm, I use them on several kubernetes clusters in the past few months and don't see any issue yet.
The main issue was ECR has a slightly different authentication model than docker swarm. The whole '--with-registry-auth' only partially works when you are using ECR. Unfortunately, it works just enough that you think it's working, until all your tokens time out and a worker can suddenly no longer pull an image.
Our common failure case was an image becoming unhealthy or a node being drained. When that image would try to be restarted on a different worker, if that worker did not have the image it would try to get it from the registry. If the tokens were expired it would fail.
The only "fix" we ever found was to setup a cron job that forcibly deployed a new version of a "replicated globally" image every X minutes (where X was based on ECR token expiration). It kind of worked, but we still had occasional failures we could not identify.
I wish it worked better, because it was nice to use ECR. Frankly token expiration sounds much more secure too, but without direct support for token refresh inside the docker engine it's just hard to get everything to work
With ECR you pay for image storage: $0.09 per GB after the first 1 GB which is free
pull from docker hub once, push to ECR. then pull from ECR as much as wish
It's not transparent though.
you can set it up in less 10 min and the only thing required is to add '--insecure-registry' in your client. It is not a issue if all your machine are in private network.
still wonder how to do it in minutes.
Expect to spend 1-2 hours first time you try it until you can setup the correct DNS records, API keys and configuration.
Afterwards it's pretty hands off, every three months you'll receive an email from letsencrypt and you'll have to rerun this script to regenerate your certificates. Takes 2-3 minutes max (but of course you still need to distribute your certificates to all relevant services...)
It isn't about the money, it's about the Mommy-May-I up and down the chain with emails and meetings and careful explanations to skeptical glares. It's a psychological and institutional barrier.
> https://en.wikipedia.org/wiki/John_Thain
"Sorry I got caught. I will work harder to hide next time."
Or has toxic micromanaged structure. I've had friends who have worked places that would barf over ongoing $60/year software charges, where anything like that would have to go up to C levels and require justification. Luckily never worked at one myself, dodged that particular bullet.
Handling each purchase and documenting it in case we are ever audited requires easily $25-50 of people's time.
The headaches to spend $5 cost way more than $5.
this is where miracle of enterprise sales happen - the $5 subscription can be sold as a $50K+ deal by smooth enterprise sales who will provide the C-exec with the experience making him feel like he did something smart and great for the company.
Engineers can head this off by prepping a total-cost-to-execute analysis for mgmt. This doesn't need to be complicated. It's just some estimates of what's needed and why, and what the alternatives cost. My eng VP used to ask me for these when I'd send up a request. He wanted to know that we thought about total-cost-to-execute. He'd usually only read the exec summary and approve. If these requests are really going that far up the mgmt chain either someone isn't doing their job, or higher level mgmt are micromanagers.
I want to speed forward 5 years and see how well this ages. It reminds me of all the other comments about "of course Facebook will never require an account to use your VR headset"...
If it was $5 for access that would be a completely different situation, but a big free tier followed by $5 for unlimited is fine.
I do sympathize with Docker though: Storage and bandwidth at that scale isn’t cheap and they need to monetize somehow.
The base costs for a lot of tech stuff (like bandwidth) are so cheap. But you would never know between these bullshit "what is it worth to you" pricing models and the number of middlemen trying to stick their hand into the pot. It's disgusting.
Note could, not will
I'd even welcome much more agressive limits than what they're proposing; the current culture regarding builds and CI in general is horrifyingly ineffficient, wasteful and in the end just plain slow.
I'm looking forward to developers adjusting their workflows (and caches, etc.) to actual, reasonable limits, not just perusing the service as if it were an unlimited cost-free cornucopia of software.
I used to be in charge of the website for a company you’ve heard of. We once realized some huge proportion of our traffic originated from a hosted CI company requesting the site thousands and thousands of times (guessing one for each build they hosted) every 5 minutes.
I can’t remember what proportion of traffic it was but I’m pretty sure it was a majority, maybe even more than 80%.
When a machine issues a ``docker build`` command, the program reads the relevant dockerfile to check for any base images that need to be pulled (a la "FROM:")
These base images are identified based on the image repository, image name, and image tag. The first thing docker does is it checks its local registry and tries to find a match for the base image the docker build is requesting. If a matching image is located in the local registry, it uses that one in lieu of downloading the image.
This is significant - if your organization only uses a few dozen base images from DockerHub, those images will only be downloaded by each build node _once_, then never again.
Many docker users erroneously believe that if their Dockerfile requests a "latest" tagged image, docker build will always download the newest version of the image. However, the "latest" tag is literally just a tag, it doesn't have any special functionality built in. If the docker build command finds an image tagged "latest" in the local registry, it stops there.
The only way to get docker build to always use the "actual latest" version of the base image is to add the "--pull" parameter to the docker build command. This arg will tell docker build to check the repository remote to see if the SHA hash of the image tagged "latest" has changed, and if so, re-download and use it. In the absolute worst case, this means each build node will pull 1 copy of each base image when the base image is updated. So unless you use 200 different base images that all have updates deployed to Dockerhub each and every day, you are fine.
While I agree that this is the way it's supposed to work, I have unfortunately worked at companies with "stateless" build/CI servers that download the Docker image each build.
Couldn't they remain stateless but be redirected through a caching proxy? Memoization is not contrary to statelessness.
Working for the same sized companies for a while has apparently dulled my senses. At a certain size, the capital that matters is the political capital it takes to get a vendor agreement in place to begin with. The monthly costs of the system are something you only feel through pushback on how big the repo gets, or the rate of traffic (experiencing the latter now with a browser testing SaaS)
docker run -d -p 6000:5000 \
-e REGISTRY_PROXY_REMOTEURL=https://registry-1.docker.io \
--restart always \
--name registry registry:2
That's it. Now fetch docker images from the IP that command is running on. Taken from gitlab: https://docs.gitlab.com/runner/install/registry_and_cache_se...https://docs.docker.com/registry/recipes/mirror/#configure-t...
https://docs.docker.com/registry/recipes/mirror/#what-if-the...
> When a pull is attempted with a tag, the Registry checks the remote to ensure if it has the latest version of the requested content. Otherwise, it fetches and caches the latest content.
Unless you’re using something like AWS CodeBuild that spins up a Linux/Windows container for your build environment, executes bash commands in a yaml file, and then terminates it when it is done. Nothing is stored locally after the build is finished.
I’m sure there are other similar services. Wouldn’t Azure Devops using hosted builds do basically the same thing? I haven’t used it since they changed the name from Visual Studio Team Services.
https://docs.docker.com/registry/recipes/mirror/
https://hackernoon.com/mirror-cache-dockerhub-locally-for-sp...
https://stackoverflow.com/questions/32531048/docker-pull-thr...
https://docs.docker.com/registry/configuration/
https://www.google.com/amp/s/ops.tips/amp/gists/aws-s3-priva...
If using Alpine, looks like docker-registry is the needed package and /usr/bin/docker-registry serve /etc/docker-registry/config.yml is the command line. Next to last link has information on the config file.
> What if the content changes on the Hub?
> When a pull is attempted with a tag, the Registry checks the remote to ensure if it has the latest version of the requested content. Otherwise, it fetches and caches the latest content.
If that causes a manifest pull, it counts as a pull and will be rate limited. Yikes! This could lead to wildly nondeterministic behavior.
> When a pull is attempted with a tag, the Registry checks the remote
Checking the remote is a manifest pull.
FROM foo
RUN apt-get upgrade etc
?Then why not have a set of base images, derived directly from upstream that get built every so often and have your private images be derived from that? This will not only relieve the stres on DockerHub and prevent you from having to pay the 5/month, but also give your security people a hook to run their tests and make your private images build faster, since all the system updates won't happen every time you change the code.
> Docker defines pull rate limits as the number of manifest requests to Docker Hub.
> For example, if you already have the image, the Docker Engine client will issue a manifest request, realize it has all of the referenced layers based on the returned manifest, and stop. ... <excluded> ... So an image pull is actually one or two manifest requests,
This still implies that even if you are appropriately re-using layers on your machine, with a free plan you can only do maximum 200 builds (since docker still needs to verify it has the image) per 6 hours?
This change also seems to imply that builds steps which previously did not handle/require authentication against Docker hub (it was only pulling public images, and pushing elsewhere) will now be required to auth against docker hub in order to double the number of pulls/checks/builds it is allowed?
If it does, then this will definitely be a reason to riot. It will effectively mean that anyone who wants to do more than 200 builds every 6 hours using the "right" way will have to get a docker pro subscription.
> There is a small tradeoff – if you pull an image you already have, this is still counted even if you don’t download the layers.
I expect we're just going to see a lot more recycling of build nodes once it has "used up it's docker credits".
Everybody else is living the pipe dream where they have externalised their risk and probably deserve the Docker treatment.
RedHat’s patches to Docker make this possible but Docker has refused to upstream it.
Why? All your need to do is to use your domain name when referencing the image.
Granted, there's only a couple base images involved, so CI pipelines will need updating to be more efficient in terms of `docker build --pull` usage.
And as pointed out below, even if you are intelligently caching layers, manifest requests count as a pull. As far as I know, no caching proxies exist for Docker that support limiting manifest pulls.
I guess one could have docker containers that actually run docker, but I don't see a reason to do that...
There is, effectively, no "virtualization" layer here. There are some things that if needed can cause overhead... such as the bridge networking (really shouldn't be a bottleneck for majority of people), and the CoW filesystem... which docker won't be (or shouldn't be) running on top of since, for example, overlayfs on top of overlayfs is not supported.
There is also nothing stripped down about the daemon inside of the container.
There are 2 components for docker: the daemon and the tool used to send commands to the daemon. In order for said tool to be able to send commands to the daemon, it needs a way to communicate with the daemon. Mounting the socket in the container is the easiest method.
I have a "tooling" image that consists of a set of scripts (python code) to do various things ops related. One of the things is to build new images when required. I have a script that given a git commit will detect the images that need to be build and build them. Having my tooling code in a container makes it easier to deploy and use new versions of the tooling code. I don't need anything on the host apart docker itself. No build scripts, no python.
As I said, i could be running the docker daemon inside the container, but that breaks one of my rules related to containers: containers are not virtual machines, they should only run 1 process and the output of that process should be std out.
At the end he describes mounting the socket. The tooling image which has all the dependencies needed to build will also have the docker cli installed, which is what I'm assuming you are doing.
I might just use this. Cheers!
You're assuming that the set of build nodes is relatively static.
Plenty of architectures set up autoscaling for the underlying nodes, that terminate servers that aren't being used and relatively soon enough (tens of minutes, hours) spin up new servers to replace them as needed.
Rarely do the machine images used to spin up new servers include the base images of the containers that will be spun up to replace them. Much more often, the base machine image is a base OS image, and container images are downloaded on-the-fly as needed. Essentially, the engineering cost of making image-launching more efficient was externalized onto an external provider willing to pay the price.
That is far from trivial.
Only if your build nodes have unlimited storage. If the build nodes are spun up on demand or have housecleaning tasks to prevent Tragedy of the Commons disk exhaustion, this is not true.
On the other hand, this is what caching proxies/registries are for.
Yes, it's really that simple. All those "container wars" and "orchestration wars" are a distraction from the core issue, which is that all those container and orchestration tools are open-source, and it's very hard to build a viable business on top of them. Docker tried and failed, like most startups involved.
Even here, when commercial projects are show, there is always an endless thread of free beer open source alternatives.
I would actually prefer if they made an incompatible replacement. Docker's CLI is pretty bad in my opinion.
I want to use Docker the same way I use a headless virtual machine running an SSH server. I want starting/exiting containers to be independent from their 'main process'. I want to attach/detach whenever I need to and execute arbitrary processes.
-- Just use /bin/bash as the main process
This seems to be the workaround, but I always have problems with containers exiting when I don't want them to and it's just harder than what it needs to be. I've spent a total of like 6 hours learning Docker and I still don't know exactly how to achieve this simple workflow without my containers quitting on me or attach/detach issues. With VirtualBox I can do this easily. Am I too stupid to use Docker?
-- Then just use VirtualBox
That's what I do, but I would like not to have the overhead of a vm.
To handle spurious interrupts from /bin/bash you can put a small script as the entrypoint containing a while true loop with a sleep infinity in it.
If you’re running systemd anyway, check out systemd-nspawn. Your ssh command becomes `machinectl shell user@container`. It’s a more VM-like way of managing containers, without Docker’s image distribution features or philosophy that containers should be ephemeral.
But if you want to not deal with attach/detach, perhaps `docker exec` is what you want. It doesn't affect the main process (unless of course your command you run kills the main process).
Isn't "docker exec -ti container-id /arbitrary/command" enough for that?
Kinda... It doesn't support caching layers for example which makes it very different in practice.
Except for Zoom, of course.
I'll repeat what I wrote there [1]:
If people really think this is a problem, they'd contribute a non-abusive solution. Writing cron jobs to pull periodically in order to artificially reset the timer is abusive.
Non-abusive solutions include:
- extending docker to introduce reproducible image builds
- extending docker push and pull to allow discovery from different sources that use different protocols like IPFS, TahoeLAFS, or filesharing hosts
I'm sure you can come up with more solutions that don't abuse the goodwill of people.
-----------------------------
Additionally, hosting a local network docker repo would mitigate this rate limit completely. Or straight up pay. It's not that difficult. Getting mad about a free, open-source service becoming pay to use... I couldn't imagine the gall and conceitedness.
This is a great idea in concept, but in practice very challenging.
RUN curl https://www.random.org/integers/?num=1&min=1&max=99999
Docker will cache this after the first invocation. The build is not reproducible. Now what?
Replace "curl random.org" with "nondeterministic and really expensive code build/model training/etc operation".
> extending docker push and pull to allow discovery from different sources that use different protocols like IPFS, TahoeLAFS, or filesharing hosts
This is great, if you can solve the image integrity/trust issues therein, which should be just some signing/merkle tree work.
The problem is, docker, the company behind the repo, has no control over what Open Source Joe and Developer Suzy are committing and the other developers pulling down their images. They can send out all these notices and announcements, and I think the typical reach of such things probably gets like what, %0.05 percent of the developers it needs to?
And of those, are any willing to rewrite the image to be smaller?
Or people will just add a proxy/imagestream in between instead of directly pulling from docker hub.
You can build a Docker-compatible image from a Guix or Nix package. You never have to use Docker or Docker Hub.
The limitation of Docker is that the nice semi-reproducible sandbox you get exists on top of an operating system that was not designed for it, resulting in giant blobs to get it to work. It's inefficient and a band-aid stopgap until we get to the future where the operating system is a pure function (which can be versioned in a tree, diffed, reverted, etc. just like git). If you used NixOS, you wouldn't need Docker. Sure, it's available and you can use it, but you wouldn't need to.
Most of us download our OS via torrents only, so we may as well download the images too if there was support for it.
I mean, I only do it to stick it to the people who claim torrents can only be used for piracy; I think most people prefer the simplicity of direct downloads though…
Afaik Windows is also using this to install updates where it shares the download with others in the region (1) using p2p (though i may be wrong since i don't use Windows anymore)
(1) https://www.itproportal.com/amp/news/how-to-stop-windows-10-...
In fact, there already exist several implementations of it for Docker![0, 1, 2]
Google Cloud uses an adjacent feature called binary authorization. When turned on, only images that are signed by a given authority (usually your ci/cd instruments) can be run inside your Kubernetes cluster.
Binary authorization may be a good starting point for someone trying to make bittorrent distributed images a usable thing.
That's exactly what Bittorrent does with its hash tree. You'd get the root hash (extremely tiny) from Docker Hub, and the rest of the metadata, as well as the data blocks, from the swarm. The authenticity is all handled by the TLS that serves you the root infohash from Docker Hub. It's a Merkle tree: the root hash is for the metadata, which is a list of hashes of the blocks.
You get the hash of the final result from the trusted server and hash is checked. Because of this you will never get an invalid image.
There are also some clever tricks to make sure no one can force you to start over from scratch by sending wrong data. But that's more of a detail.
When the backdoor is detected, you now need a revocation system so the distribution of the malicious image will die. You can theoretically do this on the tracker level, but people may build other trackers that may not propagate the changes.
- There is no rate limiting for paid accounts or companies from what I see in the article.
For your own company, you'd host a swarm that is firewalled in. Then when someone says "I want image xyz" the first thing you do is look for seeders in the swarm for that file. If non exists, then you initiate an http download from docker to get the image.
Now you've got fast distribution with low external network traffic.
Not sure how this would play with Cloud provider pricing, though. I don't believe AWS would be too happy seeing their services turned into BT swarms :)
1. Only some companies work like that.
2. I'd expect it to work like a webtorrent; try to download by p2p, but if that fails then fall back to HTTP.
Docker will be limiting manifest operations, not actual blob transmissions.
https://github.com/moby/moby/issues/1988
https://github.com/moby/moby/issues/4324
Or we're still stonewalling?
APK and APT caches often come hand-in-hand for this sort of thing as the logical next step is to add some OS packages to the pulled image. This also benefits from local caching and also means frustrating cache setup, certificate setup, etc.
To maximize local caching there's a lot of manual work to setup a house-of-cards series of proxies that only work on the network. Setting it up on a laptop then traveling means everything breaks in not-so-obvious ways when you leave the network.
AWS, GCP, Azure, DigitalOcean and even GitHub / GitLab all have private container registry offerings.
If your stack is on X provider, chances are you're going to use their private registry service instead of using the Docker Hub because you've gone all-in with that provider. That means private repos alone isn't enough to get folks to pay for Docker Hub.
- Every user gets 600 pulls per month for free.
- Pre/post pay per pull or buy a rate plan. Something like $100USD === 10,000 pulls on "pay-as-you-go" and prepay could reduce the cost per pull.
Otherwise, I don't think it's practical to have docker itself as a commercial product.
E.g. make base images for programming languages, systems, etc and guarantee security.
I think many small companies would like it a lot better if they got could externalize the cost of running docker images with all their dependencies.
I'd be pretty thrilled if Docker encouraged buildpack use. It would be a huge win for them.
Completely understandable imho.
kubectl get pods --all-namespaces -ojson | jq -r '.items[].spec | .containers[] // [] += .initContainers[] // [] | .image' | sort | uniq | less
... and look for anything that's not from a private registry that you control.
Give people two months to migrate? What a nightmare.
They could just inject their own TLS certificates into their VMs and then intercept Docker requests for images with their own cache.
Badly-behaved CI is a serious issue at scale.
I'm not sure if there's a better way to monetize the Docker Hub, but this seems so hostile to adoption.
Sounds like an entirely reasonable thing to start imagining.
Reliability, safety, determinism and predictability are not thrust upon someone from the commons.
I frankly find it someone atrocious and abusive that downstream systems do not adequately cache these assets. The main archive repositories should be the source of truth, but they also don't need to be the fountain.
It's so unbelievably user hostile that it seems like the result will be people just stop using Docker. The right solution for Docker is probably to spin off or monetize the Hub in a different way.
Imagine if GitHub started charging users for git cloning too many packfiles per hour.
Or you can monetize something else that's correlated to those costs to subsidize the main use case and keep new user acquisition frictionless. Like NPM charging large businesses with special needs with NPM Enterprise, or GitHub with teams and CI/CD features and GitHub Enterprise, and so on.
Docker sold off Docker Enterprise. Now they're trying to extract rents from people who are just trying to "docker build", a command whose ease of use is what drove people to use Docker in the first place.
Foot, meet gun.
Docker is breaking their main selling point. "docker build" and "docker run" should _just work_. By breaking that expectation in subtle ways they risk alienating users. By advertising that they're willing to break their main use case, they're going to alienate businesses and early adopter developers like myself.
Now I'm looking for alternative registries and making sure that devops code I manage doesn't depend on Docker. That's surely not what they wanted, right?
>We’ve been getting questions from customers and the community regarding container image layers. We are not counting image layers as part of the pull rate limits. Because we are limiting on manifest requests,
> For example, roughly 30% of all downloads on Hub come from only 1% of our anonymous users.
The limits appear to be 100 pulls per 6 hour time frame per ip address for anon users and twice as much for authenticated users. The least favorable reading of this is to assume a rolling period so lets roll with that. Pun intended.
According to docker whom I imagine is in a better position to evaluate the situation this will impact almost no users. Logically properly caching downloads seems like it would improve local performance. Do you really need, given the possibility to cache downloads, to pull a new image every 1.8 minutes in environments where paying $5 a month for individuals or $25 a month for a team would be prohibitive?
I'm going to assume that orgs manage a variety of one off and recurring expenses. I don't see how this is any different.
What I think is the most salient point is that docker is not a new endeavor. They already have many users. Acquiring new people who consume resources and pay nothing isn't a valuable proposition for them. Why would it be? Do you wish you had more roommates living with you eating your food and paying nothing towards the rent?
Given that even a "docker build" does a manifest pull, it's not just "new images", but existing ones as well.
Now expand that to a team that say, builds 10 images in parallel using Docker Compose. Now it's a build every ~18 minutes will hit the limit. Larger builds, like say a CI system building every time there's a push to a branch? Yikes.
I've expressed my thoughts in detail here about why I think they should find another avenue to monetize Docker: https://twitter.com/AaronFriel/status/1297988737981247488
No developer starting out is going to hit 200 images in 6 hours, and even if the did, they would go "heh" and then either take a break or pony up the 5$.
5$! It's way too little money for those limits, there should be brackets all the way up to 5000$ per month. Same for Rubygems and NPM. It's ridiculous that those are struggling organizations that can barely afford to have professionals work on them, when they're absolutely essential to whole industries.
I wish they would force us to pay them.
First, you have unlimited, then you have reasonably large limits. The rational is "limits are needed to reduce abuses and should not impact normal users".
Then, it starts to be mandatory to be authenticated. Again, officially for reducing abuses.
Once everyone has an account, and are used to limitations, free limits are reduced again, little by little, and finally to the point where you need to take the "pro" offer to have a normal usage.
At some level this seems to me like using my IDE and after 6 hours it would stop working or finding that my CDNJS references to bootstrap stopped working after 6 hours of my site being up. I think it is exactly as if NPM or PIP were to stop working for the day if you included "too many" packages.
I don't really have a good feeling for how these new limits might impact setting up new dev/test environments so I'll likely switch from using specific images to using generic images to limit my exposure.
For example, at the moment, if I want a new dev container for a project in python, I'll FROM python:latest and supplement with pulls from support service containers like nginx:latest and postgres:latest.
Moving forward, it would seem a safer approach will be to use pull a single 18.04 image and run the required installs into them. This super bums me out though as it seems to bypass some of the nicer aspects of getting up and running with docker.
I can imagine that this affects those sorts of operations.
That said, if your CI needs more, it's probably time to invest in a paid account if you can, and certainly if you're commercial.
`export ENGINE_REGISTRY_MIRROR=https://mirror.mysite.com`
Expect this to effect CI systems
We might decide on another, non dockerhub open registry though, and use that instead.
But it's on you to use it. (And it doesn't solve the metadata queries the way I understand it)
It's not hard to run a registry yourself. https://docs.docker.com/registry/deploying/ Registry supports all sorts of storage backends (files, S3, GCS ...). It's just a stateless HTTP server. It's not hard to configure for production (it comes with decent defaults). If you deploy this to something like Google Cloud Run, you can get ~free hosting (sans the storage costs) for the "registry" itself plus TLS and autoscaling too (I should probably write a tutorial on this).
That means if I build something from one organization, it's going to impact the ability to build something for another organization.
This effort seems to me like they are trying to disguise a last ditch effort to stay afloat.
Now that github, gitlab, and other great image repository options exist that don't lack these inherent security/integration features, anyone impacted can easily switch providers for no cost.
This new rate limiting won't help anything for docker, it's just going to kick off the exodus away from docker. Ironically, it's this account/security/integration stuff that they lacked focus on that lost them so much financial opportunity to begin with.
I see that Docker doesn't actually offer an AWS-style enterprise account that one can use to hand authorization to developers without requiring those developers to make individual accounts.
It feels pretty sassy of docker to give everyone 2 months to shove credentials everywhere when docker themselves haven't done the minimum to make enterprise accounts realistic. Instead, they're adopting the github model of "oh, just ask everyone to make personal accounts and then include their personal accounts in the org team". That has problems.
Firstly, it puts employers in the unpleasant position of attempting to compel employees to make legal agreements with third parties (docker, in this case). The correct way to do this is AWS-style, where the org itself makes /one/ agreement and then delegates that agreement via access keys. This is the minimum I expect from enterprise account systems, hard fail for docker.
Secondly, it's a clusterfuck to manage. You end up with an org filled with random-arse account names that you can't really audit, and you don't know who has access to what. If employees leave the org, it's hard to ensure that their access is revoked because the access takes place entirely outside the standard account domains.
Github has recently improved this a shade by adding ADFS authorization to org accounts, but that involves asking employees to tie their personal (and all github and docker accounts /are/ personal) account to their work ADFS account, which is a shitty half-solution.
All things considered, docker made this problem for themselves. They've spent /years/ working hard to get everyone to make docker accounts and push everything to docker hub instead of fostering an ecosystem of registries by different orgs for different purposes. All of a sudden it's now "too expensive" and they're dropping the hammer on everyone to sign up and push credentials everywhere with very little warning, whilst not doing their half of the work by making a proper delegated authority account system.
Doesn't fill me with confidence for their future as a stable platform on which to base a business.
All the things docker has been working for to enhance the build tooling will now be more difficult to use, even if the user never stores a single image of their own on docker hub.
One thing this is bound to do is to make the process of using docker a bit more complex. Explicit registries will probably start to be used everywhere, which is something I welcome. But it seems like a really poor decision by docker, the company to do this: they’re going to drive people off using docker hub.
Yes, docker is a struggling company which sold some lines of business and is now trying to reinvent itself again towards developers. Given other recent moves, I'm not sure the new leadership understands how to do this.
They also probably made a bad assumption in that they could define the only container format before standards bodies got involved. Blitz Scaling was the mantra of the time, seems to have bitten those back who took a bite of that cake.
best because it's hermetic
worst because it's wasteful and isn't great at any kind of graph-based step caching (even with buildkit graphs, that's not going to integrate well with your package management and build tool)