I also completely lost the plot here...
AFAICT, the thread went something like this:
The top-level concern is something like this: professors use their trusted relationship to schools in order to make bank on expert witness fees, which feels a bit corrupt and calls into question the researcher's motives.
A rebuttal to this concern is that we can side-step that issue entirely because these data sets should be public anyways (anonymized, of course!). This obviates the above concern, since the researchers won't need to compromise themselves in order to get exclusive access to data that allows them to be expert witnesses and rake in $$$$.
But the problem with that proposal is re-identification: if we can't make the data anonymous, then we all agree that it shouldn't be released (implicit in the "anonymized, of course!" caveat to "just release all the data" proposal).
Then you pointed out that even for more important data like healthcare data, FDA apparently has ways of allowing release of data that takes into account the risk of re-identification risk (I didn't know this; thanks for sharing!)
Then dragonwriter and you got deep into the weeds on HIPPA stuff.
TBH I have no idea which of you is most correct here. But anyways, there are two ways for this conversation to go:
1. You are correct, good enough anonymization is possible: Stanford researchers should not be silenced; it is problematic that they have access to data other people cannot access, but the correct solution is to negate the originally problematic distinction between those researchers and the general public by making data public. Then there is no reason for the researchers to agree to these contract clauses, because they will have access to the data.
2. dragonwriter is correct, good enough anonymization is not possible: We can go back up to the top-level concern and observe that "just release all the data with anonymization" isn't a feasible solution to this problem. Or maybe there isn't actually a problem here at all. IDK. But in any case, "obviate the problem in the top-level post by releasing anonymized data" isn't a workable solution.
Again, not following closely enough to have an opinion, but that's where we are now.
I think a good compromise position is that we should have a law stating that K12 data should be available to certain education researchers -- subject to IRB approval and so on -- without any other strings attached. Including "don't sue me" clauses in releases of public data sets does feel like an inappropriate abuse of student privacy concerns.