Another pattern that might work, which I didn't include in the article, is to scale-out the modular application (assuming it's 'stateless') into several clusters. Each node in the clusters still has the whole modular application. However, each cluster will be responsible for handling a certain part of the API. Then, put a load-balancer/API gateway in front that can route different functional parts of your API to different clusters. Scale up the individual clusters as required by load. Even though all nodes contain all modules, depending on which cluster they're in, only a certain subset of modules really takes up CPU cycles. There's still no node-to-node communication necessary, since all nodes contain all logic.
Certainly not a pattern that's always applicable, but I've used it with success several times for webapps with REST backends.