Having som spaghetti code is a natural consequence of this exploratory, iterative process. Applying good SW engineering practices to this exploratory endeavour is just a natural consequence of the process when you don’t know if your code will be of any use before you finish running and checking the test results. Why would you bother modularising and doing test coverage on something that is very likely to be thrown away after one or 2 runs?
I say this from 28 years of experience both as a data engineer and data scientist. I am a good python developer, and I can write production grade code. But I won’t refactor my code into that until I know that is the code that generates the right model. And I certainly can’t specify this particular code before writing some dirty version of it, testing it and confirming that the model it trains passes some statistical tests, at least.
Basically, the code is not the product - the product is the result of applying some transformations on data and running that through some ML or statistical algorithm to generate a model. Transformations and algorithm being unknown to be useful until tested, hence specification being unknown until coded, run and tested.
Having one person write user stories and a second person, who understands the domain less than the first person write the actual code would be twice as much work for a worse result.