Sure, I’ve done the same thing in the past; I used to use print debugging all the time. I still do, if I can’t use rr or Pernosco.
But I think that you’re underestimating the time that it would take you. Sure, if you knew ahead of time what you would need to log, it might take an hour or even less. In practice though you’re going to end up going through that loop dozens of times, adding some logging to both sets of source code (or removing something that turned out to be useless or confusing), rebuilding, rerunning the test case, and diffing the outputs.
Much better to record everything, and I do mean everything, using rr. Then you can add and remove data from your queries until you understand the bug (or rather, one of the handful of bugs that all existed at the same time) and how to fix it. No changes to the source code needed, no need to recompile them or rerun the test case.
However, I am perfectly willing to suppose that you could do in hours what would have taken me days. There are a lot of people in the world, and some of them are bound to be better engineers than I am. In that case, I suspect that using Pernosco, you could have done it in minutes instead of hours. I think it’s safe enough to say that gives me a straight 5× improvement to my abilities, and it would do the same for you.
Incidentally, roc has talked about adding automatic diffing of recordings to Pernosco. This would be diffing between multiple runs of the same program, rather than between two related programs that happen to be trying to accomplish the same task, but I can well imagine how much time it will save. Imagine those intermittent tests that fail one time out of a hundred. Currently you can rerun the test until you manage to capture the failure in a recording, and that helps to debug the problem quite a lot. But then imagine comparing a successful run to the failed run, so that you know the bug is in one of the differences between them. It’ll give anyone superpowers.