A Container Is a Function Call
glyph.twistedmatrix.com
glyph.twistedmatrix.com
https://github.com/gijzelaerr/kliko
Here is an example kliko file for a container defining the input and output:
https://github.com/kernsuite-docker/lwimager/blob/master/kli...
Also it feels like you are just making a seralization format on top of Stdin/Stdout.
The assumption is that you have containers that operate on input files and generates output files. The behavior of the container depends on the given parameters which are defined in the kliko file. The user (or runner) will supply these parameters at runtime.
To illustrate, we use it for creating pipelines in radio astronomy where we operate on datasets of gigabyes or bigger. most of these tools are file based, they read files in and write files out. It is all quite complex and old software, so Docker is ideal for encapsulating this complexity. A scientist can easily recombine the various containers and play with the arguments. By the split of input/output the container effectively become 'functional', no side effects and the results can be cached if the parameters are the same. The intermediate temporary volumes can be memory based to speed things up. We use stdout for logging.
I grabbed the repo and clicked through a few links in the docs and hit a 404. I searched on google and found a link to a way of running it just locally simply but that doesn't work with the new version. Then I followed the instructions and hit a problem installing something to do with k8 about mapped paths and the fix printed in the console doesn't work.
I understand that this is a personal complaint and others might not care at all about having it setup locally because it solves the big problems so well but I just want to try it at least locally.
I think this article nails some of the basic issues people have with docker. That is, your app is usually not entirely self-contained. One generally needs to have a db, a cache db, some networking, something else, and dash of CI.
To run a containerized app ecosystem, you either type a load of crap in the `docker run` statement or you annotate the heck out of the compose file/dockerfile. Either way, you pray it works... I know that I have to wade into logs and docker exec bin/sh because of some kind of silly-ness a lot more than I would like.
The article is correctly, imo, pointing out there should be a better way to wire up basic containerized ecosystems.
But if the Docker image is a service... We have a set of terminology (and technology!) already for services -- let's use those, rather than continue to force the analogy all the way down.
Why can't my app's ORM function call the DB container and do stuff? Why do I have to wire them up twice (code and OS layer) or more?
Maybe containers are like objects. Objects have state and multiple methods.
I personally think the object analog is a for more useful. Rather than reinventing your own container composition system, you could simply adopt a dependency injection approach. It would be nice to see containers-object publish interfaces in a language/transport independent way that layers interfaces, endpoints and transport mechanisms.
name: vanilla
summary: Vanilla is an open-source, pluggable, themeable, multi-lingual forum.
maintainer: Your Name <your@email.tld>
description: |
Vanilla is designed to deploy and grow small communities to scale.
This charm deploys Vanilla Forums as outlined by the Vanilla Forums installation guide.
tags:
- social
provides:
website:
interface: http
requires:
database:
interface: mysql
https://jujucharms.com> $ juju deploy mediawiki
> $ juju add-relation haproxy mediawiki
> $ juju deploy mariadb
> $ juju add-relation mediawiki mariadb
> $ juju deploy memcached
> $ juju add-relation mediawiki memcached
> $ juju expose haproxy
That's really cool! Why isn't this talked about more? Bias against Canonical?
It takes some steps towards resolving some of the issues in the article as well as a number of other headaches I've encountered when trying to bend docker to my will and build a development environment I can use every day for everything without having to mess around with basic plumbing.
I can post a show HN about it in the days ahead if there's any interest - it's not anywhere near where I want it to be (least of all in terms of code quality) but the amount of time I can devote to coding is about to drop from "near zero" to "really REALLY near zero" for a while, and it's very usable and handy so it might be worth just tossing it out there as-is and coming back to it at a later date.
If we talk about purely functional configuration and runtime changes those seem to be like the projects to focus on.
Nix/guix solves some cool things, but doesn't do isolation (as in process isolation). It can do it (namespaces live in the kernel after all), but it's not a first-class thing.
guix environment --containerhttps://www.gnu.org/software/guix/manual/html_node/Invoking-...
Everything else can be handled outside of the container.
guix environment --ad-hoc ruby --container -- irbThe idea is a generalization of try/catch in which all effects are accomplished through handler blocks that receive a continuation. Just like when you make a kernel call, your program is a continuation given to the kernel. Usually the kernel calls you back once, but some syscalls return twice (fork) or not at all (abort) by manipulating processes.
Coupled with reusable handler block bundles, you can wrap a "container" around any expression in your entire program.
They even serve a similar purpose: During the 32 bit to 64 bit switchover, I'd occasionally install 32-bit libs into a chroot jail to more easily compile software that depended on being 32-bit, without wanting to pollute the global space with 32-bit packages.
containers as an expanded chroot jail.
Could you point me to a succinct description of exactly the semantics of containers as an expanded chroot jail? I've been looking, but have so far not found anything.I like to think of namespaces as converting certain global variables in the kernel to local variables for each process. These local variables are then inherited to child processes. chroots are the simplest example, although they predate namespaces. Ordinarily you'd think of a system as having a single root directory; somewhere in the kernel is a global variable DIR root. But in fact, each process has its own root directory pointer in the process structure. Most of those pointers have the same value in every process, but if you run chroot, you change that pointer for the current process and all its children.
The list of possible namespaces is in the clone(2) and unshare(2) manpages (`man 2 clone` and `man 2 unshare`); look for the options starting with CLONE_NEW. They all change some pointer in the process structure, either to a substructure of the original pointer (like chroot does), a deep copy of the structure, or to a new, empty structure.
CLONE_NEWIPC changes the pointer for routing System V IPCs. CLONE_NEWNET changes the pointer for the list of network device to a new structure with just a new loopback interface. CLONE_NEWNS copies the mount table instead of keeping a pointer to it, so your process can unmount filesystems without affecting the rest of the system, or vice versa. CLONE_NEWPID changes the pointer to pid 1 / the process ID table to point to yourself (effectively a chroot for process IDs). CLONE_NEWUSER changes the interpretation of user IDs, so you can have UID 0 in your process be a non-zero UID in the outside system. CLONE_NEWUTS creates a copy of the structure containing the machine's hostname, instead of keeping a pointer to the global structure.
cgroups are resource control. They let you say that a certain process tree has some maximum amount of RAM, or CPU shares, or so forth. This is useful for making containers perform the way you want, but doesn't really affect their semantics.
> Cgroups can be guaranteed a minimum number of "CPU shares" when a system is busy. This does not limit a cgroup's CPU usage if the CPUs are not busy.
Granted my english isn't the best, but this doesn't seem to indicate any throttling.
I'd guess that it's just a matter of how you view it - putting a minimum number of CPU shares on the root cgroup is the same as putting a maximum on the rest of the cgroups, right? But maybe one or the other documentation is wrong.
I'm not so sure about your guess, because that apparently works only when CPU is fully utilized. I'm not so sure about my reading comprehension either.
converting certain global variables
How do the copies of the variables that become local relate to the originals? What happens at
read/write, what happens if the host OS changes them? In other words, what is the semantics of
sharing/copying these variables?For a chroot, you get a pointer to a subdirectory of the host root directory. Changes within that directory are visible in both directions. CLONE_NEWPID and CLONE_NEWUSER work similarly; every process has a PID and a UID outside of the container (that is, in the root namespace), but a subset of PIDs and UIDs are visible in the container, with their own values. Creating a process in a PID namespace causes it to get a PID counting from 1 in that namespace, as well as a PID counting from 1 in the parent namespace. A user account in a user namespace has a value (which could be 0) in the namespace, as well as a mapped value in the parent namespace.
CLONE_NEWIPC and CLONE_NEWNET create new, empty structures for the IPC namespace and network stack. Changes in one namespace aren't visible to another. You can move network devices between namespaces by using `ip link set dev eth1 netns 1234`, which will move eth1 out of the current process's network namespace and into process 1234's namespace. (This is occasionally useful with physical devices, but more useful with virtual devices like veth and macvlan.)
CLONE_NEWNS and CLONE_NEWUTS create a deep copy of the current namespace's mount table and hostname/domainname strings, respectively. Further changes in one namespace do not affect the other.
Has that been written up somewhere in a suitably abstract form?
Is there a list of variables that are affected, and how they are affected.
In particular, I wonder about network interfaces. Say your hardware has a network interface that's got MAC address 0b:21:b5:e2:11:22 and IPv4 address 123.234.34.45, how are these addresses affeced by cloning?
Since most people don't have spare physical devices, there are a couple of approaches using virtual devices. You can create a "veth" device pair, which is basically a two-ended loopback connection. Move one end of the veth into the container, configure them as 192.168.1.1 and 2 (or whatever), and set up NAT. Or you can create a "macvlan" device, which hangs off an existing device and generates a new MAC address. Any traffic destined for the macvlan's MAC address goes to the macvlan device; any other traffic goes to the parent device. So I can move the macvlan into the container and assign it the address of 123.234.34.46, and it will ARP for that address using its own MAC address.
The container also has its own routing table, iptables (firewall) rule set, etc. And anyone listening on a broadcast interface in the host won't get packets destined to the container, or vice versa. It's basically like a virtual machine.
like a virtual machine.
That's surprising. Containers were advertised to me as being much more lightweight than conventional virtualisation (e.g. VMware, Xen), because the former, unlike the latter, share the ambient operating system. completely separate network stack
What does that mean exactly, does the container copy the actual code, or is the network stack's code shared, just run in a separate address space?Specifically: puppet, chef, salt, juju, heat and others will happily deploy your application (container or not) and provide you with a generic interface for it. Docker is just one of the tiny building blocks here. They'll either check the model before doing anything or will check if it works by trying. Types or not, the system for implementing those interfaces exists.
But every month there's another configuration / orchestration / deployment system coming out. Some will use it, but it will only worsen the fragmentation. I feel like it would either have to be built into docker and enforced at that level, or it would be yet another system that 99% of operators doesn't use.
It works well for data processing tasks - my docker images are crawlers, indexers or analytics code in python or R. Deployment is quite simple - just push docker image, and it will be picked up on the next docker run. Images can add items to global queues, and for any bigger data they write things to a shared database.
Where each container is described by what types of things it takes as inputs and what it outputs -- ie as functions.
In practice we've built hundreds of these and composed them in endless variety for a large range of genomics tasks, processing petabytes of data at Seven Bridges Genomics.
[1] https://github.com/opencontainers/runc/tree/master/libcontai...
It turns any binary into a more portable executable.
Great post!
large programs written in assembler in the 1960s included exactly this sort of documentation convention: huge front-matter comments in English prose.
That is the current state of the container ecosystem. We are at the “late ’60s assembly language” stage of orchestration development. It would be a huge technological leap forward to be able to communicate our intent structurally.
Upon reading this, my first thought is: Please don't say XML!