Drawbridge
research.microsoft.com
research.microsoft.com
The Graphene Library OS[2] is a similar implementation for Linux and was released a few months ago. In particular the Graphene Host ABI[3] is adapted mostly from Drawbridge.
[1]http://vee2014.cs.technion.ac.il/docs/VEE14-present601.pdf
[2]https://github.com/oscarlab/graphene
[3]https://github.com/oscarlab/graphene/wiki/Graphene-Host-ABI
Two notable differences:
On the one hand, seccomp provides much more flexibility about the subset of kernel API offered to the process, rather than just saying "here's 45 syscalls".
On the other hand, Drawbridge claims to run unmodified Windows applications; they may have an efficient mechanism for trapping NT "syscalls" and redirecting them to their "user-mode NT kernel" ntoskrnl.dll. However, this might just mean that they run unmodified applications making Win32 library calls, and the libraries have been modified, in which case programs making NT kernel syscalls would not run unmodified.
I'd really like to see a standard mechanism on Linux, similar to the "personality" mechanism, that augments seccomp with an efficient process for defining a new "syscall" layer. That would make sandboxing much simpler and more efficient.
[1] http://research.cs.wisc.edu/areas/os/Qual/papers/exokernel.p...
I believe this is not really necessary. The syscall ABI is not stable from Windows version to Windows version - ABI stability is instead provided via the userspace DLLs (kernel32, user32 etc.), which are the official API applications are expected to use.
Some processes might invoke the syscalls directly, but this is a narrow use case (e.g. security software or copy protection wrappers might use syscalls to bypass userspace API hooks).
"Persistent Compatibility: allowing host and application to evolve separately. Changes in the host don't break applications."
"A library OS is an operating system refactored to run as a set of libraries within the context of an application.
While Drawbridge can run many possible library OSes..."
This implies to me at least that the app and the "OS" the application sees is a complete isolation from the host OS.
But I might be misinterpreting things here, the video seems to go into a better explanation
http://channel9.msdn.com/Shows/Going+Deep/Drawbridge-An-Expe...
Strategically this could allow Microsoft to drop lots of backwards compatibility cruft in their mainstream host OS and vastly reduce development cost and complexity.
See my previous comment here https://news.ycombinator.com/item?id=8245682
Infact, Docker started as a wrapper around LXC. It literally configured and shelled out to lxc-start in order to orchestrate containers.
This however, is much much different to cgroups/namespaces.
What the Drawbridge paper describes is a full user-mode kernel . If you want the analagous implementation on Linux look at User Mode Linux, or the Graphene stuff that has already been linked to.
What does docker do, that app-v can't?
If the features 800+ syscalls can be put into 45 syscalls, why not make an operating system that has only 45 syscalls?
Then make the kernel modular.... and we've made a full circle in the OS design :)
Because much of Microsoft's licensing revenue is contingent upon continuing to support the edge cases that are inevitably not part of the set of programs that can be dropped into such a sandbox without problems.
[1]http://arstechnica.com/information-technology/2014/04/the-in...
NT is sometimes mistaken for a microkernel partly because of the graphics driver problem, and partly because it contains a module which Microsoft actually refers to as "the microkernel". This part consists mostly of the scheduler.
What is true about the NT kernel is that it's modular, with strict internal API separation between each subsystem, and that kind of design was (as far as I know) inspired by actual microkernels, as was the idea of hiding the kernel behind "OS personalities" such as Win32.
While NT is not a microkernel, from a kernel API level it appears to basically be one, only missing the fact that everything is not actually isolated.
WinRT is COM based, similar in concept to Ext-VOS, the percursor of .NET, before it got folded into .NET.
http://blogs.msdn.com/b/dsyme/archive/2012/07/05/more-c-net-...
There is no docker like solution for Windows. All the big players (VMware, Microsoft, Symantec, etc) do tricks to isolate the applications. The tricks are instrumenting API calls and adding filtering drivers. With these solutions only less than 70% can be virtualized and the process can be really difficult.
Because from what I've seen of the current state of docker, non trivial linux applications seem to have issues in docker as well because they depend on specific things which are not being namespaced well (say /sys manipulations for example, ioctls, or even use filesystem specific APIs).
Docker seems to work well as long as one stays close to web server functionality. (aka LAMP like stacks which tend to only manipulate network sockets and traditional files).
One difference in the approach is that, for example, with Docker.io you can have your own isolate network interface while with the current Windows approach this is not possible.
At this point I doubt Wine would be interested, unless somehow Microsoft released most of it as OSS.
I might have to try it to understand it, but it is exciting.