For example a modular system is not going to help you when a module is misbehaving since the whole thing still runs in the same process (JVM in our case). If someone introduces a resource leak in some trivial module the whole thing still comes down.
For example a modular system is not going to help you when a module is misbehaving since the whole thing still runs in the same process (JVM in our case). If someone introduces a resource leak in some trivial module the whole thing still comes down.
Another pattern that might work, which I didn't include in the article, is to scale-out the modular application (assuming it's 'stateless') into several clusters. Each node in the clusters still has the whole modular application. However, each cluster will be responsible for handling a certain part of the API. Then, put a load-balancer/API gateway in front that can route different functional parts of your API to different clusters. Scale up the individual clusters as required by load. Even though all nodes contain all modules, depending on which cluster they're in, only a certain subset of modules really takes up CPU cycles. There's still no node-to-node communication necessary, since all nodes contain all logic.
Certainly not a pattern that's always applicable, but I've used it with success several times for webapps with REST backends.
As for OSGi, OSGi is complex; I dare to say that the complexity of OSGi rivals that of a microservice setup. If I had a nickle for every classloader issue I debugged... ORM was especially fun (For example wrote this piece way back: http://www.datanucleus.org/products/datanucleus/jdo/osgi.htm... ). But I must admit that I don't think we could have created (and maintained) such a large modular application without OSGi. Debugging itself is also way more complex. When an issue arises you spent a lot more time tracking down which module is misbehaving. Even though we had inserted lot's of probes (which ended up in graphite) and log statements (which ended up in graylog) to counter that.
In my experience writing smaller, simpler applications (which I acknowledge also have their own complexity with distributed debugging) are still easier to understand then an modularized application.
I once worked on an e-commerce system where we deployed many copies of a monolith. Most were serving user requests. Two were running scheduled jobs. One was something to do with a distributed cache (this was a while ago). One was processing events off a queue. Having one codebase and one binary made development and releasing easy. Having multiple roles made production simpler to understand and more reliable (i think).
Another time, i worked on business web app. Again, one codebase, one binary, three instances serving users, two instances serving a data export API, and some number running scheduled jobs. When the API went down - which it often did - users were not affected. That company also tried splitting services out of that monolith; one was pretty successful, but one was a chronic failure, because so many features involved coordinated changes to both codebases.
Modular or micro-service make no real difference here, except that a bigger collection of modules in a single JVM will take a bit longer to startup, but if you really can't periodically restart your service to reclaim lost resources due to 24/7 requirements, you will need to have several JVM instances serving the same service anyway.
This makes start-up times mostly irrelevant since any decent load balancer will handle seamless handover almost trivially.
Microservices is a rather resource expensive solution to modules bringing the JVM down for whatever reason.
Usually the right thing to do is to decrease amount of state in the application, think in microservices without HTTP and within the JVM, and finally: Run a a bunch of small JVM's instead of one large.
The upside of this approach is that when/if you need to migrate some modules to actual microservices, you have already made part of the work, already tested that you can run multiple instances of the same service, and all without the significantly increased operational and development complexity of microservices. Microservices is good where it really is the only option - eg. heterogeneous technology stacks and subsystems with different release schedules.
Simulating the grid/web of microservices including network errors like timeouts, latency and dropped connections in a petri-net (or similar) simulation tool can be a rather eye opening experience if the set of services isn't trivial. The set of weird behaviours a network of services can exhibit is staggering, even without a single line of code.
Toolwise, there are the ProM tools (www.promtools.org), which although they are mostly for mining process networks from logs can also be used for simulation and analysis. A bit of a warning though: It's a rather strange piece of software, not very stable (for me), but it has a ton of modules related to process mining and simulation.
Lots of the petri-net stuff available is generally a bit dated, as although it's a useful, and quite simple formalism, it only ever appears to have caught on in somewhat niche areas.
I agree that tracking down problematic modules is difficult. And OOM Exceptions (due to leaks) are actually the only serious problem where we run into problems during runtime. On the other hand, even in these cases its often enough to just restart the single OSGi service (if you know which one).
For NeoSCADA it was definitely the right choice. We have servers running uninterrupted for months (if not years). And to be able to hotpatch a library on the fly, is really really great (although seldomly used).