Once you have found a product-market fit and what you're offering has stabilized to a large degree, then it makes total sense to go back and rewrite things to improve performance and eliminate technical debt.
Once you have found a product-market fit and what you're offering has stabilized to a large degree, then it makes total sense to go back and rewrite things to improve performance and eliminate technical debt.
Ignoring performance is a gamble and in the worst case one might find that improving performance means rewriting everything.
In my experience, if you set things up properly the only part of the server side stack that is unavoidably difficult to transition is the underlying data management layer.
If you have good examples (outside of latency constrained applications like multi-player gaming and finance) where a decoupled micro-service architecture would need to be completely rewritten (rather than just rewriting hot services/aggressively caching), I'd love to hear them.
What do you do if you wrote your app in Python and you need high-performance multi-threading? What do you do if you chose Java and your heap is way too big? Get ready for a lot of pain, that's what.
As for the problems you specified, usually only a portion of your application work flow will require multi-threading or big mem support, in most non-latency constrained applications those portions of the work flow can easily be split off into independent services, which can be rewritten to be performant.
Obviously there are client-only applications (for example, a 3D modelling package) where it makes sense to go straight to the high-performance solution, but even in that case there are exceptions (for example, PyMol).