Docker 0.11 Released
github.com
github.com
https://github.com/dotcloud/docker/issues/928
I was surprised to see that nothing had been done about this; a comment there mentioned that someone was able to clean up 19GB of space once he figured out what was going on. The solution is easily google-able since so many have run into it, but cleanup of unused resources should be a first-class operation in every virtualization solution.
http://stackoverflow.com/questions/17236796/how-to-remove-ol...
The other suggestion, put forward by a docker dev in another comment, is to have a clean command for unnamed/untagged images. Our build system has 'buildbot' (python package compilation) and 'software' (test environment) images, so this sort of automated cleaning would get rid of all of that extra crap that we have lying around.
The big issue for me is stacks of changes. Like git, anything you commit, remove something from, then commit again still takes up that space forever until you consolidate down. Which you can't do automatically.
We don't plan on "keeping it that way". Disk clutter after lots of builds is a problem which we're very much addressing. It's specifically a top-level docker clean command that I'm not a fan of, because it doesn't capture a clear intent - what are you cleaning, and who determines what should be kept and what shouldn't? In general I like to avoid Docker having to guess what you want - the proverbial "magic".
A few notes:
* The reason Docker can clutter your disk is because we made an explicit tradeoff: it's better to keep too much data than to remove too much. That is why we're able to say "yes, you can run a database on Docker. Neither the container nor its data will be removed, at any point, unless you explicitly request it". Everytime I hear someone complain "there's too much stuff on my disk", I rejoice that I am not instead hearing "Docker ate my data".
* We have already introduced a few improvements. For example Docker used to keep the intermediary container for each build step (for later inspection of the build output, transient state etc). Now, by default, we don't.
* The most elegant fix, I believe, is to use naming and tagging as an indication of interest. If you are interested in keeping an image, give it a name (for example with 'docker build -t'). When an image no longer has a name referencing it, instead of keeping it around as an anonymous image like it does today, Docker can remove garbage collect it. That should take care of the problem elegantly and without requiring a new, clunky top-level command.
TLDR: the "disk clutter" problem is the result of a conscious tradeoff which we are glad we made. We have started fixing it, and there are more fixes to come.
> We don't plan on "keeping it that way".
Cool! Can you point me to where the discussion on this is occurring? I was confused because the issue I referenced above was closed without a reference to any issue open on how this problem will be addressed.I have a docker-cleanup script, which removes untagged images and containers: https://github.com/blueyed/dotfiles/blob/master/usr/bin/dock...
docker clean remove-untagged
docker clean remove-old 4w
It's bizarre that the Docker team is so oblivious to the issues running Docker in a real world environment.
If it were entirely up to me I would never call it 1.0, because part of me wants it to be flawless before anybody uses it in production. But it's not up to me - a lot of people are already using Docker in production, and there are many organizations already in process of offering commercial support and services for it. Clearly they consider Docker acceptable for their needs at some level - and at some point we'll need to acknowledge that by releasing 1.0, no matter how flawed I find it to be. The improvements won't stop at 1.0 :)
I would be pretty happy if Docker just had a simple implementation of /dev/log by default, so containers could rely on it being there. Then you could easily redirect everything to a production grade syslogd.
My biggest outstanding issue with Docker (Linux really) is the storage system. None of the three options make me feel like I'll be able to sleep at night. Aufs seems broken[1] and will never make it into the kernel. DeviceMapper seems okay, but I'm not sure how well tested the Docker usage is or how it performs. Btrfs seems ideal, but again, that's still an "experimental" file system.
I was even toying with the idea of implementing a Docker storage driver using rsync (rsnapshot?), for simplicity/stability, but not sure if that would even be possible.
Obviously something that will inevitably get resolved, but my biggest stumbling block so far.
(comment copied from the other item about Docker today)