> I'm not sure what you mean about it being too risky to add regression testing later?
Sure, let me expand on that. It's uncontroversial that code has to be factored for testing, projects don't usually end up easily testable by accident. You can't necessarily write a testing harness around a project as a black box. How repeatable is its output at all, how easily can you add new inputs and expected outputs, how long does it take to run, how many other dependencies does it have and how easy are they to isolate, etc.
Some types of projects are more amenable to this than others. In general, if it is a CLI with some input files and some output files, and the output is always deterministic from the input, then regression testing is as simple as collecting some representative inputs and outputs. Services with stateful APIs and multiple backend dependencies are just about the opposite of this, and require a lot of engineering to harness at all.
I have inherited a number of HTTP and RPC services where the outputs were not repeatable at all, for a mix of good and bad reasons. For example, generating random IDs as part of the output sounds like "best practices", but it makes it much more difficult to test a chain of related requests (where the IDs have to be consistent throughout) and assert their responses. At best, it requires a lot of test driver code which is itself hard to maintain. Even including timestamps can interfere with this, though that's generally easier to mock out than random IDs are.
That's to say nothing of how many projects think it's fine to create a bunch of global state, often including global "helper" threads that can't be explicitly controlled by anything else. These either have no effect on testing and thus aren't actually tested in any meaningful way, or they have strictly bad effects and make things less deterministic and hygienic.
So if you want to refactor a project to be more deterministic, have fewer and simpler dependency edges to mock out, have no global state or background threads, etc. then that can add up to a huge refactor that, in itself, risks subtle regressions which you don't yet have regression tests to catch.
That's what keeps happening. You can't have thorough regression testing until it's hygienic and deterministic, but refactoring it to be hygienic and deterministic risks regressions that you aren't ready to test for.
Breaking this cycle requires creating deterministic regression tests that you can't yet run against your existing project because it's not deterministic yet. One way to do this is by manually capturing a bunch of real examples (using real data, to cover the edge cases that synthetic tests miss), making a deterministic version of each, and making that a test case for the new version of the code. Then you rewrite or refactor the code until it's deterministic enough to pass these tests. If the existing implementation is a tangled mess, and it usually is, a rewrite can easily be more practical than a refactor.
I contend that it's not even enough to make it pass each case individually. To prove that the logic is actually hygienic, you have to engineer it to pass all cases as if they were real parallel transactions in a single instance of the system, which also limits the test case design because they have to be independent enough in that dimension while still being inter-dependent enough to prove the corresponding logic works.
I've done all of that and more for a single complex backend with lots of mutation APIs, many having pages of business logic, operating on over a dozen different entity types with subtly different storage- and API-facing schemas, using real requests and responses running in parallel with real threads. It's a huge investment, but after building the test driver itself, all code added to the project has been extremely easy to test and not a single regression has made it past this testing. As much work as this was, it took less work than some individual small refactorings did to the original project, and now it just pays off every day forever. This is not an isolated example, but it's my best and most recent.
> It sounds like one of your first steps in the rewrite is building such a test suite for "any use cases you can reproduce from the old project"
Yeah, but one of my points it that you can't even run this against the old project if it's not deterministic and hygienic yet. It's a set of potential test cases for a future refactor/rewrite that's more testable. Once you start down this road, you have to push through to completion, because every change made to the untestable old project is another way the test cases can fall out of sync and risk the new project diverging.
> but in my experience it's hard to make the rewrite-or-not decision until you have that. Only after the investigation to understand the features and behavior do you have enough info to know if you can fix it in-place or not.
I totally agree, and that's why this might only be possible a year or more into the project. I don't think anybody should attempt this when they first take over a project. I do think people should be free to make this call once they're familiar with a project, and leadership outright banning rewrites out of principle cannot possibly help with that.