Open Container Initiative specifications are 1.0
coreos.com
coreos.com
I'm very excited for all of the ideas we want to work on now that this milestone is out of the way, and we have a solid base to improve upon. But of course we should have a moment's rest to appreciate how far we've come.
Also, here's the official press release from LF: https://www.opencontainers.org/blog/2017/07/19/oci-v1-0-brin...
I'm really digging the end result of the spec - it mostly defines a bunch of conventions on top of easy formats. Kind of wish the marketing around it would be less buzz wordy, because the end result is quite straight forward.
* (It's usually a tarball; nothing stopping you from sharing images over torrent for example.)
I applaud this effort because Docker is sloppy, but I also realize that this isn't usually how technological adoption occurs. Usually the sloppy implementation is standardized, e.g. JavaScript, POSIX, CSV, Markdown, Ruby and Python to some extent -- rather than something designed from scratch or designed with fewer implementation artifacts.
Maybe C++ is an exception, where things are actually standardized and there are multiple, healthy, competing implementations.
https://github.com/opencontainers/image-spec/blob/master/imp...
https://github.com/opencontainers/runtime-spec/blob/master/i...
Now that things are 1.0 we expect many people to start implementing the spec. For example, the transition steps required to add OCI v1.0 support on-top of Docker v2.2 support for an application primarily involves changing a few strings. See the implementation notes here.
https://github.com/opencontainers/image-spec/blob/6060c3b6c6...
Containerd itself is not at all Docker-specific, you can use it independently for anything that requires transferring and running OCI containers. Because it's bundled with Docker, it's by far the most widely installed OCI implementation, easily in the tens of millions of nodes.
We also released an oci-image builder that uses a different model from docker build: https://github.com/oracle/smith
I see you missed the 1990 and early 2000's part of C++ history.
Turbo C++, Borland C++, Watcom C++, Green Hills C++, Microsoft C++ 7.0, djgpp for MS-DOS.
Turbo C++, Borland C++, Watcom C++, Green Hills C++, Microsoft C++ 7.0, C++ Builder, Symantec C++, Metrowerks C++, C++ Builder, Visual C++ for Windows 3.x onwards.
MPW C++, Metrowerks C++, Visual C++ for Mac OS.
Borland C++, CSet++, Visual Age for C++ for OS/2.
aC++ for HP-UX, xlC++ for Aix, Sun Forte C++ for Solaris.
gcc for most UNIX systems.
These are just the ones I remember without having to do a search.
Writing portable C++ code in those days was quite an adventure, specially since we only had the CFront followed by C++ ARM as standards, while ANSI/ISO were working on the first actual standard.
It was and is a standard that MANY independent implementations tried to conform to!
That's what this is. OCI v1.0 is pretty much "what Docker does".
Docker does use the OCI runtime-spec internally now (with plans to support the image-spec at some point), but that happened after the spec had become fairly stable.
The most important thing is not where the spec design came from, but that it's now being maintained by a community that will make sure you don't end up with fragmentation and lock-in.
While Docker is mostly OCI compatible these days (I say mostly as they extend their tools beyond it, making it hard to replicate the build without using docker itself. Not in-and-of itself a bad thing, but still annoying for portability), it shouldn't be confused with the OCI itself, which is there to help safeguard against companies like Docker going off and doing incompatible things.
runC is the most commonly used - it is used by containerd (which is docker's wrapper), garden (which is cloudfoundry's wrapper), and CRI-O (which is kubernete's wrapper).
rkt is another container company, whose wrapper is (afaik) just called rkt, and whose runC equivalent is systemd-nspawn, which does not meet OCI spec last time I checked.
maybe i have the wording of this wrong, rkt people please advise
ive spent the last week learning about containers via rkt, having deliberately avoided the docker hype for the last few years.
ive mostly ported my setup to rkt, using acis because thats what all the rkt docs talk about. ive invested time into the 3rd party dgr build tool, which works really well and is based on acis, and cant be easily ported to oci. ive set up CI, deployment etc. now it turns out ive misunderstood the situation and im going to have to throw it all away at some point soon.
this is basically why i didnt bother picking up docker before. or is it moby now? yeah.
should have just stuck with openvz, vagrant and shell scripts, theyve worked fine for the best part of a decade.
i was worried i was going to go back and look at the docs and see numerous glaring advisories that I'd ignored, having opted instead to read blatantly out of date pages, but no, apart from a few vague parenthetical warnings the docs all talk about building and working with ACIs and the appc spec.
As long as images share layers, network traffic would be the same in either scenario.
Reduced transfers will be achieved by formats that have a more granular filesystem representation. Right now, we are really held back by the continued use of tar, but that is another problem.
Having more granular layers can help a bit but in practice it hits its limits very quickly because higher layers will necessarily get swept up and have to be rebuilt when layers below them that they don't actually depend on get rebuilt. This rules out having a core image and a set of modules that each downstream image may or may not need (to be mixed in or not per-image). Thus, when you use the layer-based system in practice you usually end up with very little layer sharing between images, outside of the core Linux distro layers.
As far as metadata is concerned, your suggestion is exactly what OCI does. The problem with ACI is that it embeds a large part of the metadata into a tar file, which has to be fetched in its entirety. OCI is mostly metadata scaffolding, made up of indexes, manifests and configs that can all be fetched without large bandwidth requirements.
The compositional aspect that you've brought up has been explored and it doesn't make a whole lot of sense to cram that into images. Typically, such a system requires composing container filesystems through named references, allowing components to be independently rebuilt. Because this composition often relies on details of the target deployment media (orchestration system, container runtime, specific operating system, etc.), especially in how it deals with the security of name resolution, baking it into an image format leads to massive inflexibility.
Moving it up the stack works much better and avoid the technical complexities that come with doing it within images. We can see this image composition in action in k8s PODs, docker stacks and other systems. Such constructions can be distributed through OCI images, through media types, but they are based on the compositional capabilities of the target system.
But out of curiosity, why do you say that a dependency graph would be better here?
In terms of on disk storage and transport of layers, the layers are already content addressed.
With a flat list, the same logical set of files will have a different checksum depending on where in the final rootfs it needs to end up, right?
Technically you can share blobs even between an image that is `FROM centos` and one that is `FROM ubuntu` as long as the layer hashes are the same.
If you have a hierarchical dependency chain, and addressing content based on the content that comes before it in the chain, you can only share blobs between images if they have a common ancestry.
One image has a large file at a specific path, added via something like: "COPY largefile /etc/path1".
Another image wants to add this same file to a different location:
"COPY largefile /etc/path2/".
My understanding is that these two blobs will have a different hash and can't be shared because the changeset includes both the file itself, and the destination location in the image.
Is that correct?
To be honest, when I jumped from PHP dev to more app-based backends years ago, I thought it was fairly crude to have a stateful server handling everything. Now everyone's starting to move back to (hosted) versions of what amounts to "bring up and tear down the entire app on each request."
PHP is even moving toward "stateful" server handling. Just look at hhvm
Lambda == PHP. It's full circle. Don't try to pretend otherwise.
Never heard of HHVM (I've been out of the PHP world for a while), but I know PHP's default operational mode is complete start up and teardown on every request.
(I'm the CNCF executive director.)
The upsides of this is simplicity, permanent 'containerization', better handling of state, automatic updates if all this is handled with a package manager and finally the potential to have more locked down containers via analyzing the dep chain.