The crux of my argument is that some projects really are costly enough to maintain that a rewrite is an overall lower cost (pricing in risk), while others are fixable in-place and a rewrite is an unnecessary cost, and it's rarely obvious which is which for all of the same reasons that project planning and cost estimation are infamously difficult.
> The important part is that any improvement needs to be broken up into small enough pieces that each can be shipped separately. You don't want half done, unshipped work to be a continuing feature of your development process putting an overhead on everything you do.
This is an ideal but it should not be a requirement. If you "need" this for any improvement, you're going to miss out on a lot of the biggest possible improvements because they have exactly the far-reaching impacts that make them harder to pull off but also worth much more when you do merge them. If you reject any such refactorings, you'll only improve in small local ways and never large global ones.
Best case, you can try to get the best of both worlds by supporting both old & new interfaces to the same improved implementation, slowly migrating old edges to new ones. That's also just an ideal, and there are many ways it can prove impractical, e.g. subtle divergence between old and new types which is useful for the new implementation but makes it harder to interoperate with the old one.
The fact is, sometimes a design has a big enough problem that a large change will pay off, and sometimes an implementation can be bad enough that this change cannot safely be made within the existing implementation. Usually when I've seen those, it's because the regression testing was so inadequate that you can't make changes with confidence, and yet the code isn't factored in a way to introduce the testing without the refactoring itself facing risk of regression, etc. and it's a total deadlock that destroys a project.
Many, many real world projects end up in this state. Often it's because leadership prioritized deadlines more than quality, and promised to "fix it once it's launched", but then nobody was confident changing something that sorta kinda worked. If you don't invest in regression testing from the start, it'll always be too risky to add later. At that point, you may as well build a new project factored for safe maintainability including its own regression testing that will pay off forever, including testing that it does not regress on any use cases you can reproduce from the old project. You have to do something like this to break the deadlock, so it may as well fix other deficiencies as well.
I have saved several mission-critical FAANG projects this way, and even my managers agreed that it was a huge success despite the general resistance to rewriting large projects. It even takes less time than people will assume, because once you factor the new project for confident maintenance without regressions, you become far more productive working on it than anyone could ever be on the old project. You get there sooner than you expect to, and it pays off more than anyone can imagine because they're so used to the problems of the old project.
I'd also like to add that while a bad project can limp along for many years, it faces a different kind of problem good managers should fear. Only a very capable engineer can maintain such a project with a low defect rate, but they do it with great stress and frustration on their end, because truly poor maintainability hurts even the best engineers. The better the engineer, the more they feel the deficiencies of the project, and the more likely they are to want to leave. An engineer like that MAY be able to pull off a rewrite, but if you ban it, they're more likely to leave than play along.