51% of 4M Docker images have critical vulnerabilities
thechief.io
thechief.io
The vast majority of those 4M docker images are stale and probably never pulled anymore.
For example I probably have a good 100-200 public docker images on my account, virtually all of them will have critical vulnerabilities as I don't maintain them. Those are images I used for past personal projects, or testing something quickly, demo projects, iterating quickly on something etc. As soon as I am done with them, I forget about them and will never invest the time to maintain them, probably no one will ever pull them either (why would you pull an undocumented `randomuser/test-ml-my-shitty-project:latest` image?) I suspect a huge share of the 4M docker images analysed fall in that category.
Artifact Hub (which lets you find Helm charts, OLM operators, etc) will show scan details when it has enough information. If we look at the operators (and these are available in openshift today and considered curated by Red Hat) we can see a lot of operators with images that have issues... https://artifacthub.io/packages/search?page=1&ts_query_web=d...
I know this doesn't directly deal with Docker Hub but does highlight the cultural problem with addressing security.
I'll say that I poked a couple of them that matter to me to let them know about the security issues. They hadn't even noticed. Ick.
* coin miners, including images designed to mine coins, eg. kannix/monero-miner
* images for "Hacking Tools"
These two categories account for 64% of the 51%.
didn’t realize crypto ignorance was still that pervasive and discriminatory
I can see “potentially harmful” to a place the image was deployed on without authorization or against terms of use, or if the image contained a miner and didnt tell the user.
Maybe if you’ve got a sympathetic manager. Otherwise you might just get a reputation as a whiny trouble maker.
As much as I love Docker for development, going all-in on containers for prod was a huge mistake. Which is becoming clearer now that RedHat pulls free CentOS images, after taking the so-called Linux community further and further away from POSIX and other portability means with systemd, namespaces, oci, and whatnot.
Do you think Java .jars in Maven repos or .NET .dlls in Nuget are faring any better?
Find solid dependencies, don't use too many of them, make sure you update them, use static analysis tools and dependency scanners, etc. Same drill as for anything.
That's hilarious simply because most heavy Docker users I've interacted with basically use it as a way to build houses of cards which glue-sticked piles and piles of dependencies together rapidly.
This is largely because application complexity in some domains and expectations of development pace have ballooned to insanity but there are plenty of heavy Docker users that do this even when these sort of pressures don't exist.
I think as software progresses and more development techniques become easier to leverage and automated (correctly and incorrectly), we're getting to a point advanced statistical software packages have been at for quite some time: no matter what data you throw at the software for analysis and what analysis you choose, the software has so much experience baked into it that you'll get some result that makes assumptions what you're doing is valid that may look like reasonable outputs even if they aren't. If you do need to do some preprocessing, much of that is even automated so it's incredibly easy to do something very wrong if you dont know what you're doing but not realize it. Automation has made it much easier to fail late instead of fail early.
Docker isn't the disease, its use in such manner is just a symptom.
"Allow CRUD of names" ok, you'll have it next sprint. "Allow upload of their resumé docs" no problem, two more weeks. "Parse the docx and PDF files users upload and extract the contents and fill our db schema with relevant data". Well shit. The specs to understand either of those formats are gigantic and now I'm forced to cobble together something using more dependencies if you want it done in less than two year's time.
If, for example, a file format is really complex, it has most likely one implementation that almost everyone uses. If there is a problem with it, you will hear about it. Use that.
On the other hand, if a file format is really simple or you only need a simple subset of it, rolling your own could avoid you getting 0-dayed when the most popular implementation has a vulnerability.
vs getting hit when your vulnerable implementation becomes popular. Not an obvious choice indeed.
Which is not going to happen, because there's no point in publishing something that is tailored to your specific purpose.
Docker is like shipping EAR + Websphere + OS in one package.
For jars/dlls you look up dependency scanners and hook them up to your CI/CD.
For Docker images you look up image scanners and hook them up to your CI/CD.
Hot-patching apps to use newer versions of JARs because of security bugs, outside of the update stream of the actual app itself, is not really a good idea due to the risk of regressions. Docker has basically codified the failure of dynamic linking in the real world.
Also, good luck doing static analysis on binary tools.
There's an entire area of work people call "ops" that specializes on managing that dependency. But they can only do that if it's dynamic and loosely coupled.
You make your own images, if you really care. You involve your "ops" people in that.
They manage the base images that are then used by your dev team. They're responsible for updating the base images.
Your dev team provides a CI/CD pipeline which the ops team can use to check that security updates don't break functionality.
It's not rocket science.
Using distroless images or "FROM scratch" with statically compiled app reduces the risks.
You still have to watch for your app dependencies updates but that's less work than for an entire distribution.
I suppose installing the dependencies ends up being easier if you know they'll end up in, e.g. a ubuntu 18.04 image, hence that's what people do.
https://hub.docker.com/_/centos
They also have free RHEL container images.
https://www.redhat.com/en/blog/introducing-red-hat-universal...
Existing docker images are still there, future images will continue being pushed (but only for CentOS Stream post-2021), and UBI exists for minor releases of RHEL which is supposed to be identical to CentOS anyways.
And considering the fact that Docker containers don't use the kernel provided by the client OS, the difference between CentOS and CentOS Stream is even smaller than it already was. The most significant changes between minor releases of RHEL / CentOS are in the kernel.
The problem is not updating your images. Running old libs on bare metal has exactly the same problem.
nothing in docker prevents updates, just people in general are poor about doing maintenance, this is a story that is old as computers, and why MS forced updates on everyone because left to their own most people choose to "postpone" them indefinitely for fear of breaking things
It sure would be helpful if the updates didn't actually break things. It's usually not the security fixes that do that—stuff rarely breaks on LTSC, or in Debian point releases—it's all the other changes that get rolled in.
There's just too much churn across the industry. Even if the changes are good, too little thought is given to the cost of change itself.
Updating libs on bare metal has maintainers and distributions, though. And doesn't force you to roll your own security infra within containers.
Even if you do a FROM Ubuntu or FROM Debian that's still a huge advantage.
If you’re using a container based on an OS, you are installing far, far less and having only one process running. That again is a huge savings: no time securing postfix and cron, etc.
Another big advantage is that you’re shifting this work from the deployment time to build time. You do it once and can test it as long as you want before deploying it to production - no more failures because Red Hat’s mirror is unavailable or malfunctioning during your prod update window.
Unfortunately, we've gotten very caught up with the boring infra parts (the containers and the orchestration of k8s). We too often forget about the other platform we should build that uses these tools.
People don't complain about mass security holes in Heroku. That's long been container based. That's an example of building a platform that enables keeping on top of security issues.
This isn't a problem of the tools we have. It's a problem because of how we use them and what we build with them.
That's because Heroku is mostly a closed platform that the average consumer does not have insight into.
You don't know what outdated dep the software that executes your dynos runs. As the case in many areas of life, ignorance is bliss.
If you want open platforms we can talk about cloud foundry. If a base image or buildpack has a vulnerability there is a process to update it and rebuild the containers for the running application without and app dev ever needing to know about or deal with it. This is considered a feature.
The idea is to look at the architecture, concerns, and how to handle different roles with different needs at different layers of the concerns.
I don't mean to hold up Heroku or Cloud Foundry as some amazing thing. Just that you can have an architecture that solves for these problems with containers.
And why not? They're shipping benefits (immediate productivity) and externalizing costs (eventual security breaches) onto the end user.
Containers are a building block that let us handle one concern. They are a great tool for some things. But, to expect the app dev concern of "just run my code" to be solved by them well is missing the other concern.
I think Docker does some things well. I think this works as intended. I just think that we need other things that use these building blocks better to create better solutions at the right level of abstraction.
At this point in computing system complexity, secure defaults over configuration is a requirement for a sane system.
As we increasingly convert infrastructure into code, this means infrastructure configuration requires secure defaults as well.
That the Docker ecosystem doesn't seem overlying concerned (as a core, drop-everything, #1 priority) about a decades-hard security problem (detecting and resolving insecure dependencies) seems shortsighted.
To put it another way, there are two realities:
(A) No tooling support is provided to address a risk efficiently and quickly = no one addresses the risk
(B) Efficient, effective tooling is built, maintained as platform evolves, and advertised to the community = everyone does what they should
I realize the structure of Docker dissuades them from actually addressing (B), but it seems like an existential threat to adoption.
How often are people who deal with OS packages working on kernel level vulnerabilities or are dealing with microcode for CPUs? If someone is working on libc or libcurl are they supposed to follow and track all of the kernel level issues for every place the code would run?
I would argue that someone working on libcurl is working at a different layer in the stack than someone working on a kernel. libcurl can run on multiple kernels.
The same idea of layers is true when you're talking about someone working on a nodejs app and those OS layers.
> To put it another way, there are two realities: > (A) No tooling support is provided to address a risk efficiently and quickly = no one addresses the risk > (B) Efficient, effective tooling is built, maintained as platform evolves, and advertised to the community = everyone does what they should
I think Kelsey Hightower said it best when he said that Kubernetes is a platform for building platforms. I would argue that container systems in general are platforms for building platforms. I know people who spend a bunch of time building a platform on top of k8s at different companies.
Yet, in these conversation we keep coming back to Docker or container platforms as being the platform that app developers _should_ interact with. App devs speaking out, people like Kelsey speaking up, and security issues like this all say that maybe that _should_ is wrong.
You don't need docker to automate your fresh install process.
Containers are a great deployment mechanism for production. Unfortunately a lot of people seem to think that they are a substitute for package management. This was obviously (and I do mean very clearly obvious in advance from first principles¹, no hindsight needed) going to become a disaster. Most people just do not care - they are either completely unaware of the problem, or they know that in a few years' time they will either have changed jobs or be promoted (in no small part thanks to their "Docker thought leadership") and it will no longer be their problem (and they would still get to claim "Docker thought leadership" success on their résumés).
¹ Also from experience - VMware came out with "virtual appliances" in the mid-2000s that promised to do a lot of what Docker would later be used for.
> By default, Prevasio Analyzer first attempts to find the “latest” tag of a container image. If the “latest” tag is missing, it picks up the last tag enlisted in the JSON file
So it's 51% of the "up to date" docker images that have critical vulnerabilities.
Whoever came up with this heuristic probably did not really understand `latest`. To be fair, though, it is a fairly misleading and confusing "feature".
I've seen headlines like this before, but at the time a lot of the vulnerabilities were in packages that were installed on the image but were not launched, or generally not exposed. I do wish for an easy way to frequently update images though (i.e. rebuilding them from scratch installing the latest packages). It's often hard to determine what you're including in your container.
Security matters, even if it is apparently about unused parts in a complex system.
If.
Story Time: At a customer I've once seen all (web) applications in their docker containerr running with root permissions. I raised an issue but the devops told me this would be fine. They simply did not care about security.
They run all their containers on two rented VPS. They dockerized everything, also their mail infrastructure, etc.. This means: If anybody found the easiest remote code execution bug in their webpages, they immediately could take over the whole fucking company. Because root in docker = root on the whole machine. Think about that twice.
Every container is one RCE vulnerability away from being compromised and escaped. if your images are distribution based, even slim ones, you’re giving the attacker a broad set of tooling out of the box.
Container runtime defaults in Docker and Kubernetes are insecure and grant attackers a lot of privilege — running as UID 0, no user namespace separation, and potentially dangerous kernel capabilities added to the container’s parent process.
That said, is it possible to disable superuser access entirely in a container? Can't login as root if there is no root.
If it's never executed, I don't know what vulnerabilities they were talking about.
Docker just doesn't use the provided interface by default. Which is a shame because a lot of users don't bother or don't know they should bother to configure it.
I think that's kind of a 'we don't care about security' move by docker and given its userbase that's a real problem.
[0] https://blog.pentesteracademy.com/abusing-sys-module-capabil...
Docker is for many a "simple" way of getting their app running, and once it's up, they never look back.
It isn't made particular easy for you. Honestly it's actually kind of a pain, wherein lies the problem. Doesn't help that some languages' official images are kind of bloated themselves, either.
Shame it doesn't support Fedora. I all be definitely checking it out.
> By default, Prevasio Analyzer first attempts to find the “latest” tag of a container image. If the “latest” tag is missing, it picks up the last tag enlisted in the JSON file
from p.11 https://prevasio.com/static/web/viewer.html?file=/static/Red...
How are people handling container maintenance?
For example, I could imagine modifying SRPM spec files to also build a container (possibly even statically linked binaries inside). Then I can vendor update, and rebuild all the containers I need from SRPM; not much more complicated than `rpmbuild postgressql`, and re-deploy the emitted postgres container.
One of the main reasons is that Heroku's "Stack" is updated monthly (sometimes more frequently if there are security issues). You don't get this when you deploy docker to Heroku.
For better or for worse, Docker images will continue to run.
"This app is still using the cedar-14 stack, for which the end-of life window has closed as of November 2nd, 2020. The cedar-14 stack will no longer receive security updates and builds will be disabled for apps running on the cedar-14 stack."
Because people with low-level technical knowledge thinks that once everything is on Docker they're immune to lot of vulnerabilities.
So they run images without hesitation.
Presumably images from reputable vendors like alpine:latest and openjdk:jre-alpine are OK?
Looks like the vast majority are images that are intended to mine coins, but there are a few "normal-looking" ones in the top as well
the rest of that 51% had vulnerabilities of some sort.
now, while I understand the value of these analysis tools, I dislike these hyperbole they put around them.
why?
many of those vulnerabilities might be in the base image that is not actually actively used. now, this isn't great, a point of docker is to limit your base environment, but people do take fatter based images in order to make their life easier.
Lets take an example, imagine you have a base image with curl installed. the application actually never uses curl or libcurl itself, but the version of curl installed has a cve against it (perhaps even a critical one). is this a "good" situation. Not really, but the application itself as provided by the docker image isn't really vulnerable in practice as it doesn't use curl.
But, all these vulnerability scanning tools have no way to determine if curl is used or not, so they just scream "security vulnerability". Where a more nuanced take would be, the application probably isn't vulnerable due to curl being installed, but it probably be better to create a leaner image.
About 5 years ago we decided to take a path less-traveled with regard to applications development and deployment. We decided that our application is humble enough to run entirely on 1 powerful server. Sqlite is more than enough for our persistence needs. We don't need a huge orchestration of machines to get the job done. Instead, we decided to use self-contained deployments of .NET Core applications with all required dependencies embedded. This means we can email a zip file to a customer, they can extract it to any fresh windows/Linux vm, and then our code bootstraps the entire show.
All of our dependencies (the ones we are legally accountable for) are easily tracked in this approach. Introspection of msbuild dependencies is far easier than understanding the full extent of a docker image's vendor scope.
So you instead created, from scratch, a way to run all of your stuff directly on the host with absolutely no isolation? Indeed, you have eliminated side channels - by making them main channels!
I like how we are able to frame basic application development like its somehow insecure by default and only with the blessing of containerization technology can it be made secure.
If your application has serious side channel considerations, you should be running it on a dedicated host (with or without hypervisor/containers/et. al.). This is precisely what all of our clients do when hosting our application in their environments.
Docker does not magically protect your application from others running on the same physical machine. The marketing materials may present it in this light, but there are foundational security considerations with running 2+ applications on the same physical CPU that go way deeper than OS-level containerization primitives.
Like a first class way to start with
FROM scratch
and a way to configure it to pull only from a local, private registry.It seems docker (the company) has actively worked against this. The redhat versions allow configuring a private registry, but they don't worry about pissing off docker.
> and a way to configure it to pull only from a local, private registry.
One can set up an air-gapped environment with access to a private container registry only.
I think they should allow people to configure docker to be local, and allow some safer or maybe less promiscuous ways to use it (and it also goes around the ubuntu firewall)
https://stackoverflow.com/questions/33054369/how-to-change-t...
which led to this:
https://github.com/moby/issues/7203
and this comment:
It turns out this is actually possible, but not using the genuine Docker CE or EE version.
You can either use Red Hat's fork of docker with the "--add-registry" flag or you can build docker from source yourself with registry/config.go modified to use your own hard-coded default registry namespace
I'm just a web dev, so it's far beyond my skill set, and it could be out there already.
Bonus if it's pluggable into CI/CD systems, I guess it'd be practically useless if it wasn't.
Any idea for multi-tenant Docker/K8s free setup?
Something which will work for the next 10 years and all small micro-services/API will be isolated.
I do love "git push dokku master" developer experience with Dokku though https://github.com/dokku but it's based on Docker.
Which way I should look at?
KVM/QEMU?
Firecracker? https://firecracker-microvm.github.io
My imagination tells me to be with tool which are inside of Linux already ... or some very well tested tool like Nginx and limit CI/CD with Github as they seems will stay in business till the end of my life ...
I dont want to be the part of someone else ecosystem. I just want to self-sufficient, self-hosted and secure. Is it too much to wish?
Kidding aside, this is the real reason I only trust images from official vendor accounts.
It was gone within two weeks of me leaving because no one could understand what the point of it was since you could just use docker.
I'm surprised it's as low as 51%.
I'm certain it has critical vulnerabilities. But it doesn't matter because it's not running anywhere.
The subject line of this post should be something like "51% of jackets found in used clothes store ineffective when worn as bulletproof vests"
Containers mixes the concerns. We now ask those concerned with application code to be concerned with the platform they're running in. To do this takes away time and energy invested in their application.
So, it's no surprise they don't put much time in to it.
The days of letting people who are concerned with the platform be concerned with this are gone in this season. We now overload the app dev with more stuff to be concerned with.
Bringing up Bazel to deal with system dependencies is going down the road of DevOps rather than Dev. Docker is pitched as Dev. A lot of them don't care about the lower level ops.
There is something to be said for separation of concerns and targeting each concern well. Our current problems are because we're doing a poor job mixing concerns or even understanding the separation.
Yes, this takes time and energy, but so does anything worthwhile.
Now, we're moving a bunch of that ops work over to be work of the app dev. That means more context switching and knowledge outside the area of most concern. All while companies what increased velocity.
Why should developers always be concerned with the platform details?
If the developers should be concerned with the platform should ops also now need to be intimately concerned with the details of the apps? How much less time could they do on ops if they have to know all about the apps?
> Why should developers always be concerned with the platform details?
Every platform has constraints. Do you have effectively unlimited memory, unlimited storage, unlimited bandwidth, zero latency and infinitely scalable CPU available for the process(es) you are responsible for? Can you be sure that given system call is available and behaves exactly the same everywhere?
Of course not. If you were to act as if these things were true at all times, you would create something with sub-par performance and reliability, eventually driving up the cost of deployments to impossible levels.
I've been working in software since I was a kid. I remember working for an ISP in the dialup days. I saw the separation of roles between those concerned with the platform and those concerned with the running applications even back then.
To call it a brief fad, I think, misses a lot of what's happened in the market and the different roles with their different concerns.
> Every platform has constraints.
We live in a world where you have a platform built on a platform. To a node.js or PHP developer you'll have people thinking of those as the platforms. They aren't thinking about Ubuntu, CentOS, Windows, or the other lower level systems.
This is why a JS dev will code something up on a Mac and then run it on Linux in production. node.js is the platform for them.
But, with containers we're also asking them to make the Linux distro with all the things related to that as part of their platform. Not to just make it work but keep up on security. This is the wrong level in the stack for them. If they put a bunch of time into being concerned with that it will slow their velocity on the node.js work and be a big context switch (along with different knowledge in this context).
It looks like I see a bunch of Ops/DevOps people saying that app devs should be concerned with their stuff in addition or in displacement to their own. That's not happening well for a reason.
Oh, wait. </sarcasm>