Why and How Java Developers Break APIs
arxiv.org
arxiv.org
Breaking changes are sometimes a necessary evil. But they don't belong on minor releases. There should really be a stronger community standard that package maintainers sit on their hands and simply deprecate calls they want to get rid of until they can sweep them up in the next major release. It's also not cute to abuse semantic versioning by incrementing the major version number several times a year. (I'm looking at you, Guava.) That sort of stuff might be fine if a packages were relatively isolated affairs, but then the rest of the ecosystem doesn't really do anyone any favors by pulling in deep hierarchies of transitive dependencies that make it very difficult to limit one's exposure to the bad players. And then the build tools get to pile on and (as Maven does, at least) default to unpredictable dependency conflict resolution rules that obfuscate any latent problems.
I suppose it's no wonder that Java's the only place I've ever seen something like package relocation. But, at the same time, that practice has got to be the most distressing code smell I've encountered in my career to date.
Hopefully Java 9 makes some of this better (we're not ready to upgrade to it, yet), but I can't shake the feeling that this stuff really shouldn't have to have been such a hassle in the first place.
I've got a couple of questions here because it's an area I find interesting.
Do you have a list of offenders (in addition to Guava) that you frequently get bitten by?
- What other ecosystems are you familiar with and how do they compare (favorably or otherwise) to Java?
- What in particular do you find unpredictable about Maven's dependency resolution?
It would walk the maven dependency tree to find these, any time if found two versions of a library (say app depends on A and B, A depends on Guava 17, B depends on 18), it would log it as an resolvable issue.
On the project I tested it against, I remember it logging something like 250-1000 issues. I gave up. I'm not even sure where the code is now. Java is supposedly statically checked, but it lacks a real linker, as anyone who has ever seen a run time exception of "XXX is not a function" can tell you.
As far as Maven being "unpredictable", how it would resolve transitive dependency conflicts prior to 2.0 was undefined behaviour, and post 2.0 it is bizarre/lazy at best. It doesn't use the newest version, and doesn't flag major version conflicts. It just topologically sorts the graph and picks the highest dependency. The logic is bizarre. I have no idea if gradle handles it better, but they both face an unsolvable problem.
Shade is a pretty terrible work around. Auto-sandboxing conflicting dependencies in their own class loaders would be the most reasonable solution. As long as return values weren't passed between libraries.
Edit to add:
And this is not a theoretical problem either, as soon as you bring in a big dependency like Spring or Spark you are already 20+ levels deep in the tree, with dozens of these sorts of library conflicts, and hundreds of broken method signatures, just waiting to step on the wrong land-mine. (Not that I bet large ecosystem is any better)
Edit: Then I cross my fingers hoping there isn't a breakage or unexpected behavior somewhere.
Sadly, these breaking changes can be mutually exclusive: e.g. a method was renamed that A & B both use, so no matter which version you pick something fails.
I remember once having to do a binary search through versions of Guava to find one between 15 and 18 that happened to be the one that was "just right".
The issue here is javailable packaging system, people need to do non-API public methods.
a big offender here are the spring boot starter meta dependencies, which are a very opinionated set of mutually compatible dependencies, but God forbid that you want to use a newer version in your own project,
e.g elasticsearch 5 vs the "officially compatible" 2.4. welcome to jar hell
Opscode and many, many community cookbooks are a pleasant surprise there. Most cookbooks I've been working with have adhered to semantic versioning except for when a mistake occured. Backports of fixes are kinda accepted, but you have to do the work yourself.
I'd guess this is because someone writing a cookbook tends to be closer to operations and knows how useful a small, safe bugfix is.
After working with many languages, it feels like Java is the most stable ecosystem yet.
The JS ecosystem breaks APIs almost daily, similar in ruby (just look at factorygirl), and other such languages.
Maybe it’s better in Haskell or C, though.
People are inconsistent with semantic versioning, so it can't be trusted. Even people who have been consistent with it in the past may behave differently in the future. I've been burned too many times and I won't even do minor version upgrades anymore unless I have 100% test coverage on the parts of my code that use the dependency.
what if the patch included updating the (patch) version of a direct dependency? then you don't have much more control over what that direct dependency would bring on