What exactly was considered wrong with the output? How did this fail?
What exactly was considered wrong with the output? How did this fail?
I'm still trying to understand this, and I worry that much of it is due to my personal failings. I think I'm much worse than most people at making headway on problems when I know that my understanding is flawed, and I rationalize this as part of my moral code as a sort of "first do no harm". But I'm scared to discard my moral instincts for sake of convenience, out of fear that I wouldn't know where to stop.
In this case, I couldn't find an interpretation of the paper that matched my interpretation of the sample code, and I couldn't find an interpretation of either that seemed plausibly correct. The flaws in each seemed so obvious to me that I had to presume that I was not interpreting either correctly, and I concluded that I must be missing something essential.
There wasn't much in the way of test structure other than eyeing up the results and declaring it to be OK. I was scared that if I inadvertently implemented something incompatible with both the paper and the code, my errors would never be caught. None of the grad students seemed to fully understand the paper, and the professor wasn't familiar with the R code. My clumsiness with the terminology of the field made it difficult for me to communicate with the professor.
I feel terrible about the whole thing, but don't know what in particular I should have done differently.
I know that Petersen and van der Laan are well-respected researchers in causality that know exactly what they are doing. E.g., Petersen has a course (I think at Berkeley) that uses causal graphs which were pioneered by Judea Pearl (see comments below). I can only second the recommendation to dig into his work.