An alternative to fat JARs
product.hubspot.com
product.hubspot.com
Wait until they find out how much faster in-process function calls are than RPCs.
Honestly, if you can't trust your devs not to make a mess in a monolith, you can't trust them not to make a distributed mess with microservices.
Both architectural patterns have their uses. Neither will save you from bad developers.
I also could be wrong about this, but I like the microservice after projects reach a certain size.
I am a big fan of a well-designed SOA, but that's not properly bloggable here in 2016.
Now, OK, the JVM ecosystem lacks a good module system at the moment. But Java 9 will fix that. There is also OSGi though I don't know much about it.
It's not a panacea, it looks like you're just pushing the complexity into another place.
And no, not panacea, just a way of thinking that I think is useful.
With a single memory space it's easy to simply go around whatever architecture you "should" be following.
Microservices don't solve balls of mud, but they can make it harder to form and easier to spot forming.
Pushing complexity onto interprocess/machine communication will just result in higher levels of spaghetti. Do you really think that same people who would construct a monolith poorly would do a better job building a distributed system instead?
I don't mean to be too dogmatic; there are legit use cases for services. I am just highly skeptical of microservices as a solution for complexity.
Now that Moore's law is running out of steam, we can reasonably expect that to become false within the next decade or so. The division between those who can merely write code and those who can properly engineer software systems will become all too apparent then.
2017 will be the year of the return of the monolith: "How We Replaced 946 Microservices With a Statically Linked 10kB Rust Binary and How It Ulimately Did Not Save Us From Our Complete Lack of Anything Resembling a Coherent Business Model".
I use a framework at Google that is microservice-based, and allows for composition of microservices into an assembly. Within an assembly, calls between services are done in-process; if you are calling a service in a different assembly, it has to make an RPC.
The mechanics of whether it's a local call or an RPC are all handled at the framework leve. This makes splitting a service into its own assembly a simple task (don't have to update any code, just configuration; this can be managed by SREs).
Sounds a lot like OTP.
In Java-land, you could also just have your client and service use the same interface. Just inject different implementations for local and remote.
Once you've gone to the trouble of automatically generating the glue between the local interface and the remote object, it becomes simple to interpose further useful things - security checks, transactions if you're using a relational database, and so on. You can even have the glue code lazily instantiate the service components, so you don't need to actually load them until a client makes a call.
You'd want a cool name for it, though. Extreme Java Binding? Enhanced Jet Bus? Excellent Joy Bringer?
The approach described in this article is the way Java was designed to work; but then everyone said you needed a "self-executing" service binary for portability. Enter the fat JAR, containers, orchestration layers, etc.
JVM application servers are certainly the way to go if you run a bunch of Java. But then you're committing to maintaining a bunch of dependencies on your application servers, so really you're pushing the tech debt down the line to your ops folks who now have to ensure dependencies are synced, maven servers running and updated properly, proxies punched for new external dependencies, etc.
Jonathan: if you're listening, have you guys looked at Maven as a source instead of an S3 upload? What percentage of your dependencies are internal vs. external? We get 10-20x space savings from public-Maven-only elision.
EDIT: we've been working on a similar technique for executable JARs as well, which shares the same general optimization numbers and approach. It's a bit trickier, since we want to preserve the same 'java -jar' syntax, so requires a bit of classloader tomfoolery, and corresponding classloader awareness in code.
For WARs: In a mvn build step:
- identify candidate dependencies to exclude from the WAR. This includes a filter step to keep internal dependencies in the WEB-INF dir.
- generate a JSON manifest of all the excluded dependencies, including their local SHAs, and store the manifest in WEB-INF.
- copy a shell script into the new Maven Archive (MAR) that knows how to reverse the process above and output a WAR
Then, in our build process, we pass around the MAR files until the moment of deploy, at which time we expand back into a WAR and deploy.
We looked at just deploying the MARs themselves, but the way our containers manage file cache space, it was easier to rehydrate them pre-deploy.
For executable JARs: again, we start in a Maven build phase, and again identify the dependencies to include vs. the dependencies to exclude and fetch later. We create a totally-new JAR that contains all the classes in the thin Maven-built jar (I believe as a referenced JAR, not as expanded classes). We then bundle all the included JARs directly into the target JAR, and write a CSV manifest identifying which external dependencies are needed at runtime. We then copy a raw .class file with no dependencies [1] into the target JAR and write a manifest that identifies that class as the Main-Class. That class is responsible for invoking 'mvn' via shell execution to load dependencies, building a classloader with the right classpath, etc., and then handing off to the original Main-Class.
[1] hence CSV instead of JSON
Treat any open source project your product depends on the same way you would anything you paid cash for - store a copy in your code repository and back it up with the rest of your source code.
- if the release of the same version was changed upstream, and you want this change to be picked up, that's not a repeatable build anyway.
- if the release was changed upstream and you don't want to pick it up, configure Archiva to fail on checksum mismatch [1].
- or as an alternative to enforcing the checksum, configure your local repo as internal instead of proxied, and handle it all yourself [2].
[1] https://archiva.apache.org/docs/2.2.1/adminguide/proxy-conne...
[2] https://archiva.apache.org/docs/2.2.1/adminguide/repositorie...
We actually still produce executable JARs with this approach. We configure the maven-jar-plugin to construct the classpath at build time and add it as a Class-Path entry to the manifest. This is a special manifest entry and on startup the JVM automatically adds these files to the classpath, so our thin JARs are runnable with java -jar assuming the dependencies are dropped in the right place. The slimfast-plugin reads the configuration of the maven-jar-plugin to make sure it's generating the right paths for the dependencies. We end up with an executable JAR and no runtime or ClassLoader funny business.
I'm not saying it wasn't a dark time, but it had its merits and drawbacks just like everything else.
Also, we use Singularity as our Mesos scheduler and it handles S3 artifacts out of the box. Previously we gave it a single S3 URL for the fat JAR, but now we just give it a list of artifacts (the app plus its dependencies) and it handles everything for us so it's not much more complexity and ended up being really easy to integrate into our deploy process.
Joking aside, fantastic idea =)
Have you received a cease and desist from SlimFast yet for violating their trademark?
[0] - http://fortune.com/disrupted-excerpt-hubspot-startup-dan-lyo...
I. Build your app as a thin JAR, with:
a. a Main-Class entry in the manifest
b. Class-Path entries in the manifest
c. a file containing a list of dependency coordinates (group, artifact, version) in META-INF
II. Write a shell script (for whatever value of 'shell' you like) which:
1. takes a Maven repo URL and a Maven coordinate (group, artifact, version)
2. downloads the JAR from the repo, extracts its dependency list, then pulls its dependencies too
3. (optionally) somehow records which JAR is the main one, say by writing a tiny shell script or a symlink pointing at it
You might be able to use a standard embedded POM instead of a dependency list, which might reduce the work a bit, but that would then require doing transitive dependency resolution at deploy time, which is probably a bad idea.Phase I is something like 10 - 20 lines of Gradle, tops. You could pack it up into a plugin easily enough. Phase II is a similar amount of shell script, maybe more.
For extra safety, add SHA3 hashes to the dependency list file, and check them when you download the dependencies.
I've described this as using a Maven repo, but that doesn't mean it has to hit Nexus or whatever; you can just put JARs in S3 in the right layout. A while ago, i was doing this by maintaining a repository locally, and just pushing the whole thing to a Bitbucket website:
https://bitbucket.org/twic/twic.bitbucket.org/src/14ac48d4c4...
You could do something similar, perhaps going from Nexus to S3.
On deploy, the simplest way to get up and running is to use the download goal of the slimfast-plugin. It reads this JSON file, downloads each dependency (using an optional cache folder), verifies the file size and checksum, and copies it to the correct relative path. The application will then start up happily with java -jar.
At HubSpot, we instead integrated this download step more transparently into our deploy process. At build time we read this JSON file and store the dependency information in our build database. Our deploy infrastructure already accepts S3 URLs and handles downloading, caching, verifying checksums, and copying to arbitrary directories so we just piggybacked on this existing functionality to have it download the application plus all of its dependencies for us.
https://bitbucket.org/twic/ensure
Turned out to be more than 20 lines, but then i did it in Java rather than Groovy.
I do the downloads from the Maven repo (which should be your internal Maven repo!), so there's no need for an upload step. Deployment is done with a shell script with a few undemanding dependencies - unzip, curl, and openssl to verify the digests. I use a TSV file rather than JSON because it's easier for the shell script to read!
My next trick will be to figure out how to get this into a Cloud Foundry buildpack ...
Oh, the name: https://ensure.com/nutrition-products
You can go just as far with http://www.capsule.io/ and the capulet to import dependencies with Maven (from a Nexus proxy that also fronts S3)
You are now deploying thin Jars and having the artifacts cached on the local host through maven~!
Here's the reasons I see that makes it worth it:
* It reduces the build time to less than 50%
* They are doing this for 1000+ java applications
* The code vs library ratio is 1:100 (roughly) - that reduces the uploads from 50-100 GB cited in the article down to 50-100 MB.
Or could it be much more sanely organized with say 15 or 20 applications?
I know nothing about the codebase but it sounds like they may have nanoservices...
Saving traffic is good in general, insofar as it speeds deploys and allows higher container density.
But for me the key was trying to get out of dependency hell. The shade plugin isn't perfect.
I work at Pivotal, we inherited the Spring team from VMWare. When I first saw Spring Boot I wondered what all the fuss was about. Then I worked with a classic Spring-with-hand-rolled-Maven app and oh boy, let me tell you, I got it.
Having someone else level the dependencies for you? Huge. I'd be interest if this tool can help Spring Boot too, though the runtime downloading of dependencies is not without problems.
I'm used to having to worry about disconnected environments (I worked on Cloud Foundry buildpacks for a while), for which Spring Boot's JARs work a treat.
In a fully connected environment this approach looks promising. If I bump into any of the Spring folks I will mention it.
Also, we use Singularity as our Mesos scheduler and it handles S3 artifacts out of the box. Previously we gave it a single S3 URL for the fat JAR, but now we just give it a list of artifacts (the app plus its dependencies) and it handles everything for us so it's not much more complexity and ended up being really easy to integrate into our deploy process.
When something goes wrong we rarely downloaded the JAR, but if you really wanted to you could download the thin JAR and use the SlimFast metadata bundled with it to download the exact dependencies it would get when deployed.
Then I send a message to a remote agent with the command to run (classpath, main class, jvm args, args, configuration, etc). The agent then downloads the jars that are not already present in its cache with a valid SHA-1 checksum then starts.
Works quite well.
The biggest problem I've found so far is simply 3rd party dependencies that aren't packaged with OSGI manifest information. Recently I specifically had an issue with Spring in this regard (yeah, yeah, don't use Spring, I know... but there was a specific reason in that case, at least initially) that would have been easy to resolve if the Spring guys still distributed OSGI'ified versions of their jars.
Still, all in all, I'm pretty happy with this approach.
It's a mess, and I fear Java 9/10's modularization approach will even make it worse (so that instead of P2 (Eclipse's Equinox' OSGi bundles) and Maven, we'll have another set ... sigh. I'm still looking for a solution that transparently translates Maven metainformation to OSGi bundle information at some higher level (for example by hooking into the OSGi container's dependency resolution logic).
Edit: apparently, such a mechanism exists (ResolverHook)
All of this said, I'm fairly new to OSGI, so I may still run into some dark corners that will turn me off. But right now, it seems to be serving the purpose well.
Maybe something like proguard could reduce deployable jar size to only the classes used.
I assume these are daily dev deploys and not production deploys.
We do frequent production deploys (of individual services, there's no such thing as deploying our entire application). To give an idea, it's a little before 1pm here and across our team there have been 180 production deploys already today.
"Combining 100,000 tiny files into a single archive is slow. Uploading this JAR to S3 at the end of the build is slow. Downloading this JAR at deploy time is slow (and can saturate the network cards on our application servers if we have a lot of concurrent deploys)."