HNHacker News
TopNewBestAskShowJobs

avsm

1,549 karma · joined September 17, 2010

anil.recoil.org hacker @ocaml faculty @cambridge computer laboratory fellow @pembroke college cambridge hacker @openbsd
submissionscomments
avsm··on Docker for Mac and Windows Is Now Generally Available and Ready for Production
This bug's been driving us mad because we can't reliably repro it on our machines at Docker, and it only happens to a small subset of users, but is very annoying when it goes trigger. It seems to be related to the OSX version involved, but there's not enough bug reports to reliably hone in on it.

The other aspect that it may be is a long-running Docker.app -- since as developers we are frequently killing and restarting the application, it could happen after a period of time. I've now got two laptops that I work on, and one of them has no Homebrew or developer tools installed outside of containers, and runs the stable version of Docker.app that's just been released. If this can trigger the bug, we will hunt it down and fix it :-) In the meanwhile, if anyone can trigger it and get a backtrace of the com.docker process, that would be most helpful. Bug reports can go on https://github.com/docker/for-mac/issues

avsm··on Docker for Mac and Windows Is Now Generally Available and Ready for Production
This seems to work fine for me:

docker run -it -v /private/etc/passwd:/etc/passwd alpine sh (not recommended for any actual use obviously)

Is there a particular case in which this failed for you? We'd appreciate a bug report on https://github.com/docker/for-mac/issues (or from the Docker for Mac GUI, just click on "Diagnose and Feedback") so we can chase down whatever issue you're having.

avsm··on Docker 1.12: Now with Built-In Orchestration
Sorry about that -- this was a bug in one of the betas where log rotation was broken. It's been fixed about 3 or 4 betas ago via a mid-week hotfix, so the latest open beta should be just fine. Hope you get a chance to try the latest series again soon!
avsm··on Docker Betas for AWS, Azure, Mac, and Windows
https://forums.docker.com/c/docker-for-mac https://forums.docker.com/c/docker-for-windows

hope that helps!

avsm··on HardCaml: Register Transfer Level Hardware Design in OCaml
I've coincidentally just ordered an Altera SoCKit to play with this more over the summer. There's a nice tutorial on getting started with HardCaml and the SoCKit at https://ocaml.io/w/HardCaml
avsm··on Improving Docker with Unikernels: Introducing HyperKit, VPNKit and DataKit
Please do let us know if it persists, to beta-feedback@docker.com or on the forums.
avsm··on Improving Docker with Unikernels: Introducing HyperKit, VPNKit and DataKit
Containerd is just one part of what Docker does. There is also image management (signing via Notary), distribution (pull/push to the Hub), orchestration (Swarm and friends), linked services (Compose), etc. You could reuse all the rest of it while running containers in FreeBSD jails as the method of containment.
avsm··on Improving Docker with Unikernels: Introducing HyperKit, VPNKit and DataKit
We're taking it in a fairly different direction from upstream xhyve and bhyve, with the integration with other kits like VPNKit, DataKit (and soon, FSKit). We wanted to be able to diverge the build system in particular, and this makes it difficult to maintain a direct fork.

However, we're committed to upstreaming patches to their relevant projects where relevant, and so the HyperKit Git repository is as close as we can make it. If it diverges a lot in the future with file renames, we'll have to figure out the Git-fu to make cherry-picks easier...

avsm··on Improving Docker with Unikernels: Introducing HyperKit, VPNKit and DataKit
It should be possible to build a simpler port for FreeBSD with the recent Docker 1.11 release. It moved the container execution to containerd (https://github.com/docker/containerd), so that's where a bunch of the Jails logic would go.

I'm quite keen to see Docker running on FreeBSD so I can use it for my CI pipelines. I'm less interested in Linux emulation to run Linux images -- I'd like Docker support on FreeBSD to run FreeBSD images first!

avsm··on Improving Docker with Unikernels: Introducing HyperKit, VPNKit and DataKit
Open sourcing our libraries is obviously just the start. We're really keen to hear about any other uses that people might have and move HyperKit in the direction of easy integration with higher-level frameworks.

For instance, direct ports of unikernel projects to run against the HyperKit kexec APIs would be really fun. Any takers for MirageOS or HalVM? :-)

We've started keeping a list of "help wanted" issues for anyone interested in getting started with hacking in this area: https://github.com/docker/hyperkit/issues?q=is%3Aissue+is%3A...

avsm··on Reason: A new interface to OCaml
It's important to realise that an LLVM backend is not a panacea -- the OCaml native code generator is highly optimised for the calling conventions and OCaml GC already. LLVM's IR doesn't quite track values at the same abstraction level as OCaml so some features such as exceptions become very expensive if mapped onto LLVM without changes to the IR.

Another major recent advance in OCaml 4.03 (released last month) is the flambda middle layer -- basically an epic inlining and allocation elision pass. See https://blogs.janestreet.com/flambda/ for more on it, but it's already giving 10-20% performance improvement on real world code.

Multicore is also coming soon -- see https://ocaml.io/w/Multicore for a summary of activity there.

avsm··on Reason: A new interface to OCaml
There's a screencast fresh off the presses on the info page at https://ocaml.io/w/Blog:News/A_new_Reason_for_OCaml

I'm finally going to switch away from my ancient nvi setup and use Atom instead! MirageOS recently moved all our libraries over to using the new PPX extension point mechanism in OCaml instead of the Camlp4 extensible grammar. This means that MirageOS libraries should be compatible with Reason out of the box -- so it'll be possible to build unikernels from a slick editor interface quite soon hopefully!

avsm··on Docker for Mac Beta Review
> I cannot believe they are using `docker.local`. This hostname will cause nothing but trouble for years to come.

We are indeed moving away from `docker.local` in Docker for Mac. There have actually been two networking modes in there since the early betas: the first one uses the OSX vmnet framework to give your container a bridged DHCP lease ('nat' mode), and the second one dynamically translates Linux container traffic into OSX socket calls ('hostnet' or VPN compatibility mode).

Try to give hostnet mode a try by selecting "VPN compatibility" from the UI. This will bind containers to `localhost` on your Mac instead of `docker.local` and also let you publish your ports to the external network. One of our design goals has been to run Docker for Mac as sandboxed as possible, and so we cannot just modify the /etc/resolv.conf to introduce new system domains such as ".dev".

We've been iterating on the networking modes in the early betas to get this right, so beta9 should hopefully strike a good balance with its defaults. It's also why we've been holding a private beta, so that we can make these kinds of changes without disrupting huge numbers of users' workflows. Your feedback as we figure it out is very much appreciated!

avsm··on Docker for Mac Beta Review
[I work on Docker for Mac]

The early betas focussed on feature completeness rather than performance for filesystem sharing. In particular, we have implemented a new "osxfs" that implements bidirectional translation between Linux and OSX filesystems, including inotify/FSEvents and uid/guid mapping between the host and the container. Getting the semantics right took a while, and all the recent betas have been steadily gaining in performance as we implement more optimisations in the data paths.

If you do spot any pathological "spinning cases" where a particular container operation appears to spiking the CPU more than it should be in, we'd like to know about it so we can fix it. Reproducible Dockerfiles on the Hub are particularly appreciated so that we can add them to the regression tests.

avsm··on Docker for Mac and Windows Beta
The new daemon (dubbed osxfs) is FUSE-based at the moment, but also provides a semantic translation layer between OSX filesystem calls and Linux kernel events. The FUSE layer can be removed in the future in favour of a direct kernel module with this architecture, if it ends up being a bottleneck (its fine right now though)
avsm··on Docker for Mac and Windows Beta
I haven't tried the Android emulator recently, but I am interested in deploying Facebook's Infer tool on our codebase (and they've got a Docker container too for it of course). So I've filed an internal bug for us to look into HAXM and figure out if it plays well with Docker for Mac/Windows. Thanks for the pointer!
avsm··on Docker for Mac and Windows Beta
This is exactly correct. We're really enjoying working with the Hypervisor.framework, VMnet.framework, and all the various hooks Apple has exposed for apps like Docker for Mac. There are some bugs in the short-term, but Apple has been steadily addressing our Radar bugs and we have workarounds in place in the Application for the most annoying ones.
avsm··on Docker for Mac and Windows Beta
Thanks, this is useful feedback. There are various workarounds in the app to prevent such things, but the purpose of the beta program is to ensure that we catch all the weird permutations that happen when using hardware virt (e.g. the Android emulator).

If anyone sees any host panics ever, we'd like to know about it (beta-feedback@docker.com) and fix it in Docker for Mac and Windows. Fixes range from hypervisor patches to simply doing launch-time detection of CPU state and refusing to run if a dangerous system condition exists.

avsm··on Docker for Mac and Windows Beta
Yes that should work fine with osxfs (the new filesystem engine). Do you have a pointer to the specific bug in vbox so we can add it to our test suite?
avsm··on Docker for Mac and Windows Beta
Yes, quite a few issues of that nature have been fixed (and we are planning to open-source the changes later in the year once we stabilise the overall application).

The bug above has been reported to Apple and they've reportedly fixed it in the latest 10.11.4 seeds, but we've put in a workaround that detects ACPI sleep events and freezes vCPUs just before going into hibernate mode. None of the beta testers have reported any sleep crashes using Docker for Mac recently, so if you do see anything of this nature please let us know.

avsm··on Docker for Mac and Windows Beta
We'd love to get your feedback on the new filesystem engine in the Docker for Mac app. It's been a ton of work to get right, and there a few corner cases in the current beta that we're squashing, but overall things "just work" for my day-to-day Linux development on my Mac using the current beta.

At this stage, pointing it to the weirdest and most wonderful filesystem stressers you can find is welcome. We'll leap on any issues you find...

avsm··on Docker for Mac and Windows Beta
[I work at Docker on the announced Mac app]

Nathan LaFreniere (the author of dlite) is awesome, and we've been exchanging tips and tricks and areas where we can collaborate. He knew exactly where to press to find bugs in our earlier betas...

avsm··on Docker for Mac and Windows Beta
Let me explain Docker for Mac in a little more detail [I work on this project at Docker].

Previously in order to run Linux containers on a Mac, you needed to install VirtualBox and have an embedded Linux virtual machine that would run the Docker containers from the Mac CLI. There would be a network endpoint on your Mac that pointed at the Linux VM, and the two worlds are quite separate.

Docker for Mac is a native MacOS X application that embeds a hypervisor (based on xhyve), a Linux distribution and filesystem and network sharing that is much more Mac native. You just drag-and-drop the Mac application to /Applications, run it, and the Docker CLI just works. The filesystem sharing maps OSX volumes seamlessly into the Linux container and remaps MacOS X UIDs into Linux ones (no more permissions problems), and the networking publishes ports to either `docker.local` or `localhost` depending on the configuration.

A lot of this only became possible in recent versions of OSX thanks to the Hypervisor.framework that has been bundled, and the hard work of mist64 who released xhyve (in turn based on bhyve in FreeBSD) that uses it. Most of the processes do not need root access and run as the user. We've also used some unikernel libaries from MirageOS to provide the filesystem and networking "semantic translation" layers between OSX and Linux. Inside the application is also the latest greatest Docker engine, and autoupdates to make it easy to keep uptodate.

Although the app only runs Linux containers at present, the Docker engine is gaining support for non-Linux containers, so expect to see updates in this space. This first beta release aims to make the use of Linux containers as happy as possible on Windows and MacOS X, so please reports any bugs or feedback to us so we can sort that out first though :)

avsm··on A better inliner for OCaml, and why it matters
The runtime aspects of the multicore runtime has been pretty stable. Most of the effort currently is going into the algebraic effects extension that is used to map direct-style concurrency into multiple (parallel) cores: https://github.com/ocamllabs/ocaml-effects
avsm··on FSCQ: A formally verified crash-proof filesystem [pdf]
There is a gap between the formal specification and reality though. We published some work (coincidentally presented right after this paper at SOSP 2015) on building a real-world mathematical model of POSIX and using it to test existing filesystems; http://sibylfs.io

Running SibylFS on Fscq showed up some limitations:

- https://github.com/mit-pdos/fscq-impl/issues (mostly basic missing POSIXisms)

- https://github.com/mit-pdos/fscq-impl/issues/2 an interesting one where the proof about truncate was incorrect in the specification and lead to very weird observable filesystem behaviour.

Copy of the SibylFS paper here for the interested; http://anil.recoil.org/papers/2015-sosp-sibylfs.pdf

Having said that, these are edges to be polished -- Fscq is an absolutely fantastic advancement in filesystems, and it shouldn't be too many years before we can have a formally verified and specified filesystem in production.

avsm··on Unikernels, meet Docker
(I'm a slacking OpenBSD developer and a unikernel hacker, so I guess I should respond :-)

Pledging is the most usable form of priv-dropping I've seen yet in any OS, since it was designed after looking at the "flow" of how a typical daemon interacts with the kernel, rather than a low-level syscall interface that the typical programmer would have no idea about. I've used many of the different privilege dropping interfaces in OpenBSD over the years, and to recall the journey:

- Separating out syslogd back in 2003 was a fun adventure in manual privsepping before it became mainstream; I think just sshd used in it base before that. This required building an automaton of how the privileged messages from syslog would work, and carefully coordinating the state machine and message passing around that. http://marc.info/?l=openbsd-cvs&m=105967566808306&w=2

At this point we had two processes, but the root process couldn't do fine-grained dropping of privileges and just had to be carefully audited by lots of people.

- systrace came along, which allowed policies to be built about which syscalls could be called. The policy language was limited, so I got really excited by the possibilities and wrote a DSL to make it easier to build systrace policies; http://anil.recoil.org/papers/sam03-secpol.pdf. In the end I gave up on using systrace since it was so brittle to unrelated changes in libc or dependent libraries changing the order of system calls and causing apps to break all the time.

- Pledge is almost the opposite of systrace. You provide a series of human-readable strings, and the OS takes care of mapping those to groups of syscalls. This makes so much more sense, since the person making the change can also update the pledge, and applications don't break! If you look at OpenBSD current, over half of the base daemons are now pledging. At no point are there insanely complex policies like SELinux, nor are there the race conditions of systrace. In fact, the most similar approach to this I've found is MacOS X entitlements, which also let the application specify what functionality it would like (as opposed to what syscalls it needs)

So, what's the relation to unikernels? Well, they are completely the opposite approach. Instead of starting with a wide kernel interface and then restricting it, unikernels start with a very narrow interface (the hypervisor) and build up higher level abstractions as they are demanded by the application.

This happens via a series of libraries, and so the programmer can choose at compile time how to weave in privilege levels and the use of hardware enforcement (such as processes).

I can see both pledge and unikernels converging in the future, specifically by a linker that would be aware of the pledges that a library needs, and mapping the appropriate hardware enforcement into the resulting unikernel. I'm not aware of anyone actually working on this, but get in touch with me if you are interested... (anil@recoil.org)

avsm··on Unikernels, meet Docker
> If the linux people finally make isolation secure I see no future for unikernels.

You're forgetting that unikernels are a library OS and not tied to a particular hypervisor at all. MirageOS code can currently be compiled to target:

- the Xen hypervisor via MiniOS, with Mirage-supplied implementations of XenStore/device drivers/TCPIP

- bare metal and the KVM hypervisor via Rump Kernel

- UNIX binaries via tuntap (which work great with Linux containers).

And future backends -- the MirageOS frontend just needs to swap out and link in the right libraries for the desired platform. And even when Linux containers get a complete isolation story, if you build applications as unikernels you can also choose to isolate kernel components that will never be covered by the current Linux container architecture (such as the TCP/IP stack).

avsm··on Why We Use OCaml
Not in a released version, but there's very active work in trunk to make it work natively:

  https://github.com/ocaml/opam/compare/master...dra27:windows-build
A recent demo I saw a few weeks ago had everything running under Cygwin and building native Windows executables...
avsm··on Irmin: Git-like distributed DB
That's sort of the entire point. The idea is that you build datastructures using this library that operate over the three-way merge, and provide their own guarantees about merge conflicts.

For instance, a weakly consistent data structure could promise never to raise a merge conflict, and therefore be safe to compose. A stronger one could raise more precise merge exceptions depending on the exact error, which ripple up to the application.

An example of an application handling this sort of merge error is the Cuekeeper TODO manager (which is a pure JavaScript Irmin app that uses HTML5 Localstorage/IndexedDB). Try opening http://test.roscidus.com/CueKeeper/ in two tabs and creating conflicting changes, and see the Irmin merge error ripple up to the UI.

avsm··on Irmin: Git-like distributed DB
See also an early paper on "mergeable persistent data structures" that shows how to do this for Irmin-based Rope and Queue data structures. http://anil.recoil.org/papers/2015-jfla-irmin.pdf
← PreviousPage 4 of 9Next →