Docker was unavailable in Ubuntu/Debian repos
github.com
github.com
I am copying it below:
<<< Hi everyone. I work at Docker.
First, my apologies for the outage. I consider our package infrastructure as critical infrastructure, both for the free and commercial versions of Docker. It's true that we offer better support for the commercial version (it's one if its features), but that should not apply to fundamental things like being able to download your packages.
The team is working on the issue and will continue to give updates here. We are taking this seriously.
Some of you pointed out that the response time and use of communication channels seem inadequate, for example the @dockerststus bot has not mentioned the issue when it was detected. I share the opinion but I don't know the full story yet; the post-mortem will tell us for sure what went wrong. At the moment the team is focusing on fixing the issue and I don't want to distract them from that.
Once the post-mortem identifies what went wrong, we will take appropriate corrective action. I suspect part of it will be better coordination between core engineers and infrastructure engineers (2 distinct groups within Docker).
Thanks and sorry again for the inconvenience. >>>
Be sure to clean your apt cache before trying again (apt-get clean && apt-get update) if you're still seeing this.
> At the moment the team is focusing on fixing the issue
> and I don't want to distract them from that.
That might be ok for feature teams, but for infrastructure tools/services, it's very frustrating for users (devs) to be kept in the dark on the progress of the fix.At work, the incident response starts with identifying Investigators (to find and fix the problem) and a Communicator (to update channel topics, send the outage email & periodic updates, field first-line questions about the incident, and to contact those most affected by the incident so they don't get surprised/try to fix it themselves). The person who starts the incident is the Coordinator, who assigns the roles, escalates if more help is needed, tries to unblock investigations, and turns facts from the investigators into status updates for the communicator.
if a service i use is down, all i want is an acknowledgement and that "we are working on it right now with high priority". i want all available resources to be fixing the problem.
my anxiety over powerlessness in relying on others during a crisis manifests in other ways, like figuring out why i'm at the mercy of this thing in the first place, and putting alternatives in place.
but during the crisis i'll just go do something else for an hour and then read the post mortem when it comes out.
furthermore, even if you somehow could divine the future, telling the customer that you think an outage will last over half a working day is going to turn an extremely shitty situation into something even worse.
I consider high quality communicators to be a significant resource, and the problem to solve is to make sure everyone affected has the right information to make their next decision to mitigate the issue.
These communicators might not know how to solve the actual technical problem, and it'd be like putting monkeys on a typewriter to tell them to "all hands in fixing a configuration issue with the web server" for example.
I'd say apply the right people to the right problems.
Sometimes the problem is even something that some users will be able to work around if they know what it is.
If you want your deployments to be independent of the outside world, design them that way!
Like tsuresh said - stuff happens. What if you internet connection went down for a long period of time. You couldn't continue working. It takes very little to setup, gives you fall over but also makes installing dependencies sooo much faster.
You should be able to do something in an emergency, even if it requires manual intervention. If you can only shrug and wait, that's bad.
Welcome to cloud computing!
Most places I've ever worked in will have local repositories, procedures and timelines for anything from Microsoft and OS updates, to development stacks and libraries.
Neither workstations nor servers get updated directly from external/vendor/open repositories - it is all managed in-house.
Slower, yes; more work, yes; but that's exactly the type of issue it's meant to prevent :)
It's a little mind-boggling to me that anyone would rely on the constant availability of a free external service that's integral to putting their product together. I handle timestamps for codesigning through free public sites, but I've also got a list of 8 different timestamp servers to use, and switch between them if there's a failure.
If you're running in production without having a APT mirror[0] in your local network something is wrong with you, no with docker apt repos.
The question is why would anyone would expect immutability after pointing their package tools at a mutable repository?
I'm not that familiar with Docker but I am of package/dep management (from deb, jars, npm, eggs etc) and you most certainly want to use a mirrored package repository (jfrog, sonatype, or whatever) for this reason and many more other reasons (bandwidth, security, control, etc).
So if you did have issues with the outage I would look into one of those mirroring tools. At the minimum it will speed up your builds.
Physician, heal thyself.
My company's actually done the same thing before (same error), by putting Cloudfront in front of our APT repo -- it cached the main packages file inappropriately, causing the checksum mismatch.
It appear to be that way. Reminds me when all of the reddit admins were stuck on a plane on the way back from a wedding [1].
Remember kids, improve your bus factor.
http://highscalability.com/blog/2013/8/26/reddit-lessons-lea...
https://en.wikipedia.org/wiki/Pacific_Southwest_Airlines_Fli...
Captain Blackadder: Baldrick, what are you doing out there?
Private Baldrick: I'm carving something on a bullet, sir.
Captain Blackadder: What are you craving?
Private Baldrick: I'm carving "Baldrick", sir.
Captain Blackadder: Why? Private Baldrick: It's part of a cunning plan, sir.
Captain Blackadder: Of course it is.
Private Baldrick: You know how they say that somewhere there's a bullet with your name on it?
Captain Blackadder: Yes?
Private Baldrick: Well I thought that if I owned the bullet with my name on it, I'll never get hit by it. Cause I'll never shoot myself...
Captain Blackadder: Oh, shame!
Private Baldrick: And the chances of there being two bullets with my name on it are very small indeed.
Captain Blackadder: Yes, it's not the only thing that is "very small indeed". Your brain for example- is brain's so minute, Baldrick, that if a hungry cannibal cracked your head open, there wouldn't be enough to cover a small water biscuit.
Clearly, we need to operate in cells, so that no one knows everybody and will have everybody over for weddings and other parties.
I believe commercial releases are downloaded from a separate infrastructure (to be confirmed).
Either way, the availability of Docker packages, free or commercial, is critical infrastructure and we should treat it as such. IMO our primary infrastructure team should have been involved, and someone should be on call for this. We'll do a post-mortem, find the root cause, and take corrective action as needed.
Apologies for the inconvenience.
In open-source we call that "thursday" :)
Where we draw the line is if people are being intimidated, bullied, insulted, or anything that even remotely resembles harassment.
Although I personally feel that some of the comments in that thread are pretty unfair and poorly informed, they don't seem to violate the social contract.
We moved off it last year. When we went to cancel our subscription the other month downgrades/cancellations were broken on the site as a known issue; had to open a support ticket. Most of the UI issues were still present along with some new ones.
If the apt repo was compromised (but the signing keys were not), this is very likely exactly the symptom that would appear.
I don't think that's correct. It would pass a checksum test and fail a signature test with a "W: GPG Error". The checksum test is not about cryptographic security, it's just about files referenced by the Packages file having the same hash that the Packages file declares them to have. You don't need any signing keys to make that happen.
So I'd guess a rational attacker would choose a bad signature. But attackers can be irrational; it doesn't prove it's not an attack. Just not my intuition.
Just like the node builds that failed this should cause you to rethink how you mirror or cache remote resources not prompt you to complain about your broken builds on a github issue page. There may be things you'll never be able to fully mirror or cache (or could just be entirely impractical) but an apt repository is definitely not one of them.
If you're running more than three machines or regularly (re)deploy VMs, it is a sign of civilization to use your mirror instead of putting your load on (often) donated resources.
It's the same stupid attitude of "hey let's outsource dependency hosting" that has led to the leftpad NPM desaster and will lead to countless more such desasters in the future.
People, mirror your dependencies locally, archive their old versions and always test what happens if the outside Internet breaks down. If your software fails to build when the NOCs uplink goes down, you've screwed up.
So as soon as you depend heavily on external sources, you should start to think about maintaining your own mirror. Software like pulp and nexus are pretty versatile, and give you a good amount of control over your upstream sources.