Capsule Shield: A Docker Alternative for the JVM
blog.paralleluniverse.co
blog.paralleluniverse.co
We designed Nomad[1], our scheduler, with this in mind. Nomad has the concept of "task drivers" and one of those task drivers is "java." A task driver provides isolation using things like cgroups automatically, without you having to containerize your application. This makes sense for applications that are already _mostly_ "containers" (in the abstract sense): Java JARs, statically linked binaries, VMs, etc.
If you compile a fat JAR you can schedule it directly with Nomad, and Nomad will handle the isolation and resource constraints. Pretty nifty!
And, I noticed Capsule talks about OSv. Nomad can also deploy "VMs" natively on top of Qemu so you can deploy that as well.
The coolest thing perhaps is that you can mix and match this stuff: you can deploy JARs alongside containers alongside VMs and Nomad handles the binpacking and resource constraints for you. And this is why I bring up this plug for us.
That JAR better be a capsule, then :) Non-capsule fat JARs don't have enough information to configure the JVM, don't support native dependencies and in too many cases simply don't work.
These days, production and dev environments are extremely heterogenous. We're already running a database written in C, a messaging queue written in a JVM language and a front-end API written in Node and all while serving our static assets from a generic web server. And we're starting to explore Go and Rust for more specialized components. Docker allows me to abstract away all the platform-specific issues and treat all my infrastructure components as if they're homogenous. Docker compose can bring up my entire dev stack. Docker swarm or any of the other orchestration services (ECS, Kubernetes, Nomad, Fleet, etc) can manage my entire production deployment.
A JVM-specific option is really a step backwards. Using that, I'd now have to go back to treating my JVM components differently from everything else. The whole point of Docker, for me, is that none of my diverse infrastructure components needs special consideration.
It's not about philosophy but pragmatism. Most organizations are mostly Java, and even if Java doesn't make up 90% of the components, it makes up 90% of the deployments because that's what the core business software uses. Such organizations get to use a a rather crude tool that takes away some of the power of their technology stack. Why not use a slimmer, faster, less-hassle tool for the majority of your deploys?
And if we want to talk philosophy, Docker really mixes two separate issues -- portable packaging and isolation. Capsule separates them. If you want, you can launch a capsule inside a Docker container. If you want, you can directly launch it in a more convenient container simply by launching it with Shield (or outside a container by running it directly). As Shield will be open-container-compatible, there's no reason why Docker and Shield shouldn't play together. And remember, Capsule requires zero additional tools on top of what developers use anyway (it's just a plain Maven Central library that you stick in your JAR with your build tool).
Also, remember that Capsule doesn't add another build step. You need to create JARs anyway, so you might as well make them capsules. You then either have the option of adding another step of building Docker images (just because you have some other Docker containers, most of them not even built by you) or not. Capsule just saves you that extra step and still lets you isolate your app in a container if you want, while making everything faster and management+logging easier. I don't think the need/wish to use containers for isolation should dictate your packaging solution.
For most companies that operate several services, having one trusted and strong way to deploy and orchestrate is way more important than an occasional load/build time when working on a new laptop / instance.
Edit: I'm not saying Capsule isn't useful, it just doesn't make any sense to say it compares to Docker, because the use-case is quite different in reality.
...quite a bold statement. Can you cite some sources?
[1]: http://www.zipcodewilmington.com/blog/why-java-skills-matter...
[2]: http://www.infoworld.com/article/2868654/it-jobs/java-develo..., http://www.infoworld.com/article/2608294/java/employers-want...
Docker enables everything to play nicely together. Capsule seems to double down on the old, monoculture way of developing. For people that want to do that, I'm sure your solution is an excellent product. But my own perspective, which may be skewed by being in the bay area, is that those people are the minority, not the majority.
I won't say what you're saying isn't real, but it is happening on a completely different scale than the size of the Java ecosystem. Even on the JVM, what you see as "plenty of Scala" doesn't even amount to 5% -- not that it matters, but it's just an example of why I think your view is skewed.
The size of the Go ecosystem is hardly 1% of that of the JVM (rough estimate), and while Node.js/Python are "a thing" (and are the only languages/platforms of those you mentioned that are used to a substantial amount outside of startups), they are still tiny compared to the use of the JVM. Even in Silicon Valley (which is in itself a very small portion of the software world), the large "webby" application companies (with the notable exception of Facebook) are JVM shops: Amazon, Google, eBay, Twitter, Netflix are either entirely JVM or predominantly JVM, and even Facebook is making growing use of the Java.
Most of the shifts you're seeing is among non-JVM technologies. It's mostly Pythonists that switch to Go and Rubyists that switch to Elixir/Node.js, etc.. The JVM was never the biggest player in the "fast application development" arena -- at the beginning there was VB, and then Perl, and then Python/Ruby and so forth. I see no signs that there is greater movement away from the JVM -- mostly among those who realize they're better served by the "fast dev" platforms even at the cost of performance -- than towards it -- by those realizing they need the best the industry can offer in terms of performance/monitorability/tooling even at the cost of slower development.
I work at a firm where we have dozens of greenfield Spring Boot projects going on.
And if we perform the stats on our public cloud, Java dominates the workload.
However for early stage web startups especially ones that use the JVM there are only really likely to be a handful of diverse technologies in play. A typical example: nginx, jruby, redis, kafka, <primary persistence tech goes here>. Now at first glance you might look at this and see 5 diverse systems and think "these guys could probably benefit from docker", and perhaps some could, but I personally think that a growing number of shops feel like it's a waste of time and resources, and here is why:
Although there are five separate systems in the example above the truth is that developers are not sitting around writing nginx config files all day, nor kafka config, nor redis config, nor Mysql config files. Those are things that you only have to setup correctly once and then you only change gradually, a few lines here and there over the lifetime of the business until you settle in on the most optimized configuration for your workload.
99.9% of your team's changes will only touch the "application layer" which in this example resides on a single runtime environment (the JVM). Teams who utilize the JVM in this way usually don't feel "pinned in" by it since after all the performance is there, the mature profiling tools are there, the library ecosystem is there, and there are diverse programming languages, frameworks, and multiple programming paradigms all of which ride atop the JVM. examples: (Java, JRuby, Scala, Clojure, Vertx.io (non-blocking io/javascript), Lift, Akka, spray.io, Play, Immutant, Wildfly, etc)
What I'm getting at is there's really no good excuse for a team who standardizes on the JVM to ever need to go outside of that paradigm unless perhaps they're writing a C extension to optimize some tight loop, but these days the JVM's optimizing bytecode compiler is getting pretty decent at doing that for you.
If 99% of your application code (the code that is updated continually) runs on the JVM then even if you have a micro-services architecture with several different programming languages in use you don't really need to lean on a containerization technology to standardize deployment because all of it is already under the same runtime environment roof anyways. From an operations standpoint all your build/deploy step really needs to do is ship jar files around over the network irregardless whether it's a clojure service, or a scala, or ruby app. And all of it can be monitored in a clean way thanks to the monitoring support built into the JVM called JMX.
To use docker in such a situation is to just waste system resources and to increase network IO (moving heavy disk images around) for no reason. Instead you can just use shell scripts or Ansible to provision an instance and then rely on the JVM's own venerable mini-containerization primitive (the jar file) as the standard way to deploy apps across your fleet.
To each his own though..
But there are a bunch of ways capsule makes this difficult, particularly in the context of a mono-repo with a bunch of internal libraries. Capsule lacks a real transitive dependency management system, so you're stuck bypassing it entirely and having your build system compute the complete graph of transitive dependencies and write it in the jar manifest.
Even worse, on every start up capsule will open all the jars that have already been cached to validate checksums, which can take many minutes if you have a lot of dependencies.
The idea is really solid. I just wish the implementation were better.
The Maven caplet uses Aether which is the same dependency-resolution library used by Maven. If you've found a problem, file an issue and we'll solve it.
> so you're stuck bypassing it entirely and having your build system compute the complete graph of transitive dependencies and write it in the jar manifest
That sure sounds like a bug. The Maven capsule should resolve transitive dependencies.
See also this bug on maven-capsule-plugin [1]. Having tried to actually implement that strategy (albeit in gradle), it's clear that capsule isn't well-adapted to this sort of thing.
[0] https://groups.google.com/forum/#!topic/capsule-user/Mjtnvwt...
[1] https://github.com/chrischristo/capsule-maven-plugin/issues/...
1) Allow a more elaborate description of dependencies 2) Improve the build-tool plugins to make generating the full dependency tree in the build tool and passing that on to Capsule more convenient.
I favor option 2, b/c the precise resolution already happens in the build tool, and seems like an easy and complete solution. We can continue this discussion on the mailing list if you like.
In any event, as you can see, we are always open to suggestions and happy to accept PRs.
So why not just use a fat jar, as you were before?
And a capsule is a fat jar. Non-capsule fat jars frequently just don't work: they can create resolution conflicts which requires shadowing and even that can fail, they don't support native dependencies, and most importantly they require a startup script as they don't set up JVM options.
And while the word "container" is used here, we mean a different kind of containers -- those that give you virtualization and app isolation at the OS level[1]. So there is no container -- like a sevlet container -- that hosts multiple applications, and no specific programming model.
[1]: https://en.wikipedia.org/wiki/Operating-system-level_virtual...
Why?
> Image management, logging and monitoring present challenges to general-purpose container solutions, but are non-issues for the JVM.
WTF? Java applications can be a pain to set up. Jars here, jars there, Maven, Ant, XML configuration, its all a big mess. If someone can do all that for me and say "here's my Docker image that just works" I would say issue solved.
> Docker images are big and contain full-blown operating systems as they are meant to run arbitrary applications: managing their archival and evolution can become a serious hassle. On the other hand JVM applications need nothing more than a JVM and a kernel
This person seems to have a serious lack of understanding of Docker. Your Docker image does not need to contain a "full-blown operating system" - it is perfectly acceptable for it to contain just the strict dependencies of the program (the JVM) you are running. On my system java links against libpthrad, libdl, libz and libc.
> Linking: Java already supports customized domain-name resolution, so Docker’s solution of modifying the container’s hosts file is unnecessary
Does he even get what linking is about? Linking is a way of sharing configuration between containers...so the mysql host:port on one container is available as a convenient environment variable or /etc/hosts on another. What has this got to do with Java's DNS custom support?
> Monitoring and Management: The JVM has its own rich monitoring and management API, JMX
And Docker is somehow preventing you from using this?
This post reeks of Java elitism and written by people who didn't even bother trying to understand the technology they are claiming is inferior or not applicable to a True Java approach.
"Docker images are big and contain full-blown operating systems as they are meant to run arbitrary applications"
because it's exactly how you are not supposed to use docker, you are supposed to include only what you need and you most definitely do not need a full operating system inside a container that runs on a full operating system.
back to Java: our base java docker images are around 70MB, see this to see how, I'm not involved in it: https://github.com/delitescere/docker-zulu
Yes, still bigger than the uberjar but that's the not point
When I say that something is higher footprint or worse performance etc., I mean that that's what you get for the same amount of work compared to your point of comparison, not that it's impossible to make it better if you work harder.
It's more productive to develop general containers that solve issues with logging etc. for all platforms, not for just one.
So the containers are still general, and you use them for all platforms and they all interoperate. I just see no sense in requiring an extra, inconvenient packaging step for a platform that already packages applications well.
Disclaimer (deep breath): I previously worked on the CF Buildpacks team, and I work for Pivotal Labs, a division of Pivotal, the company which donates the majority of engineering effort to Cloud Foundry.
Oh: and it has a secret feature - pragmatic and friendly developers.
That could be something decent. I'd be interested to know what performance benefits you get from using Capsule to manage OSv. Does anyone have a JVM-based Docker app they could use for comparison?
Note that I said Docker-compatible. One alternative to the Docker Engine itself would be Joyent's Triton, which uses Illumos rather than Linux as the kernel (while still supporting Linux binaries), so it doesn't have the security problems of Linux namespaces + cgroups.