California moves to silence Stanford researchers who got data to study education
edsource.org
edsource.org
That's an unreasonable restriction and I expect the ACLU to win this.
It also goes beyond K-12. Quite a few parents try to get their children’s college grades and run smack into this law.
How about data pertaining to your kids?
Any Joe Taxpayer doesn’t have a right to walk in and demand any data they want from a government department. That’s entirely entirely reasonable. Anonymising data isn’t nearly as easy as a passer-by with “faker.fake_name()” may think it is.
It's very possible to aggregate and anonymize (remove PII); aggregate the records by zipcode, anonymized school ID, school district, grade-level of student, age of student, educational attainment, etc.
You agree that's possible, FOIA-compliant and has been done already for decades? Like how Census data is made public (the Census also uses fuzzing to prevent reverse-engineering to individuals, esp. in tracts with small populations).
> Any Joe Taxpayer doesn’t have a right to walk in and demand any data they want from a government department.
Red herring: the CDE is not trying in good faith to define what level of aggregated data would be sufficiently anonymous; they're blanket-opposing legitimate public access to this data (even highly aggregated) via the researchers being allowed to testify in court.
> How about data pertaining to your kids?
Absolutely can and should be disclosed, in aggregate. Otherwise you have a public entity spending $128.5 billion taxpayer money that is not performing so well, violating constitutional disclosure requirements, gotten worse since 2020, lost students to homeschooling [0] and move-aways. In any case, this isn't fodder for an opinion poll, it's what the Constitution says.
[0]: https://www.montereyherald.com/2023/06/04/three-years-after-...
Why do you expect the ACLU to win?
(not arguing. Genuinely curious.)
Looking at the details, it seems that this cannot be a blanket restriction, since a judge could compell you to provide testimony. [0] At that point it would not matter what the contract said.
You can’t pitch a research project and then go rouge and do whatever with the data.
It looks like the state is interpreting that use of student data as part of the lawsuit to ve outside the scope of the prior approvals, therefore they are preventing Sean and Tom from using the data during the their testimony.
Nothing prevents the defense to subpoena the same data and have them use it for their testimony.
I am not a crazy disciple of the 1A but that seems pretty clearly to be something the government should not be able to do. Couldn't the government just slip that language into any of their FOIA agreements etc?
It would be a very different situation for a non-government actor to have the clause.
Very scummy by government.
That's not what this case is about. The California Department of Education is not claiming that Thomas Dee misused confidential data. The CDE said that their contract with Dee means that he cannot testify against them or participate in an unrelated case against them.
ACLU is often manipulative in how these stories are framed. They are an advocacy group advocating for their priorities.
If the headline was “Researchers publish confidential student data through litigation filings”, the crowd here would be pearl clutching about that. You may recall years ago when NYC released taxi data on their open data platform, that data let you basically track the movements of frequent taxi users.
The other question is… why don’t the researchers file a FOIA, which has no restrictions.
This is not a guy who shoots from the hip.
But nobody can actually say this. Instead we have to pretend like this isn't the case. Just look at math Olympiad teams. I coached one years ago. My entire team was Asian except for two alternates. One who was Russian, and the other Indian.
Yes, environment can change outcomes ... but maybe it can't change outcomes to a point where everyone is going to perform the same. Are we going to try to get everyone's 100M sprint into the same range too? People are different.
We should give every individual the same shot at opportunities but I don't think we are ever going to make Asian kids perform at the level of other kids in math or vice versa. Its not environment. Every one of us that has taught an engineering or math course knows this. Even if we don't talk about it.
> But nobody can actually say this.
“Asian kids in the US are, on average, better at math and, furthermore, this effect is stronger the fewer generations removed from immigration they are, and is in large part due to well-established general familial impacts on performance and the selective filter of immigration.”
Race is not a useful scientific guideline for any kind of scientific study. For example: there is as much biological diversity in sub Saharan Africa as the rest of the world, but racially, the best we can do is "Black", or "African". It's a useless, dated concept that we, as species, find it difficult to work past because our brains are categorical engines.
I'm as politically "leftist" as anyone you'll ever meet, but we have to be able to do better than "Asians are good at math" to make effective decisions about education, amongst other problems. This is of course impossible with the current world and thinking. Even though I know race isn't real, I still see it. It still has an impact on my day to day actions, because my stupid brain is all too happy to categorize people on how they appear.
Taking another route: to say that Asians are good at math is categorical error. The word "Asians" represents something abstract, and abstract things cannot take action. Categorical error is basically the starting point for the various "isms" like misogyny, misandry, racism, etc.
You, I and everyone else are stuck in this way of thinking because we've been so thoroughly indoctrinated into a system of white supremacy which permeates the entirety of Western culture, it isn't even noticeable, like we're in the Matrix. It persists because it's useful for keeping the power centers that benefit from it entrenched, and everyone else divided.
We can move on from it, but I think the first thing we need to do is recognize that it isn't inevitable.
Do you think Asian people would have not come to the same conclusion, even if white peoples hadn’t said so first? I tend to think it was somewhat inevitable that Chinese, Korean, and Japanese people think of themselves as having more in common with each other than with French or Mexican people.
And even in a single location, if you look back in time, you can see how people got categorized shifting. People used to insist Italians weren't white here in the US.
Tell that to prostate cancer researchers.
To say race "isn't important" is completely ignorant.
On the surface of it this sounds absurd, because (unless adopted) people do not determine their ancestry by looking at photos of themselves. I can see getting proximal affiliations wrong, confusing or missidentifying oneself as being half Italian when they're actually half Iberian, or or confusing turkic ancestry with Persian. But I don't think people are going to not know whether they are primarily of say East asian, african, or european ancestry.
>The mapping is very noisy.
>"race" as such isn't important
Sounds like quite the leap to reach the conclusion that you're trying to make.
You could easily have a genetic predisposition to prostate cancer without being a certain race, even though that "race" may have a higher propensity for that genetic trait.
Not really, everyone knows that an albino African is still an African. Physical traits are just the most visible aspect of genetics. And your second point is just explaining outliers, it doesn't say anything.
Race is entirely a social construct. You can't do a genetic test and with certainty determine someone's race. Certain genetic traits are common among what we call races, but not exclusive.
Take a look at services like 23andMe or other services, the genetic components of race are entirely based on self-reporting, that is, we call certain genes "Asian" because people who identify as Asian had those.
It's entirely tautological.
The pairing between shapes & suits will of course be tautological because the names of the suits are cultural artifacts, but the shapes would still be distinct regardless.
> Race is entirely a social construct. You can't do a genetic test and with certainty determine someone's race. Certain genetic traits are common among what we call races, but not exclusive.
This seems confusing and contradictory. If traits are common in certain groups and not in others (needn’t be exclusive), then by Bayes rule these traits should identify groups with high probability (especially when combining multiple traits)
I'll end with a question - competitive cycling requires the same abilities as marathon running, that is, optimal oxygen intake and usage, and great endurance in the buttock, leg, and foot muscles. Why are Ethiopians not dominant there?
Also, to present a possible answer to your rhetorical question: cycling is a sport for rich people. A good bicycle, that a kid that wants to pursue the sport must get early in their life, is very expensive.
The most important factor in either sport in any case isn't the muscles, it's oxygen intake and use efficiency.
As far as your argument, by effect of selection, doesn't that meant that top runners are more likely to come from poorer countries? You also don't really need a good bicycle to train, just to compete, but even that's pretty expensive so I take your point.
If race isn’t real, then as a white person, my bone marrow should be just as compatible for transplant at the same probability for blacks, asians, and “mixed race” people as it is for other whites, right?
I totally agree that monitoring every single human being on the planet, and recording and analyzing their individual DNA and second by second logs of their biomarkers and external environment from womb to death would definitely be “ideal”. But we aren’t there. Yet.
In the context of trying to manage finite resources and time, broad messy abstractions have been and will continue to be crucially important, despite not being pure. Trying to erase things like race with an ideological handwave is harmful.
edit: also, note, you've said a lot of stuff about categories. But categories can blend into each other and cause ambiguity at the margins. But that doesn't remove the validity of there being categories. When I look at the kids who come in to competitive math programs, and these are kids who are more than a standard deviations above the mean in performance, I see a lot of uniformity. One can try to construct various explanations for this, but you can't tell me that there were NO kids from underrepresented ethnicities that had two professional parents and good exposure to mathematics early. We have plenty of racial diversity in the early math programs. And in the Bay Area there are plenty of professionals sending their kids to these programs from all ethnicities. And still, ten years later its the Chinese and Taiwanese American kids that are in the Olympiad team. And even at a lower level, say SAT math, we see the performance skewed by race in the same way. I am ethnically Indian. And Indians are as interested in math as Chinese. We all send our kids to math tutoring. We are mostly engineers in the Bay Area. And even so, at the very top of the distribution, there are some Indian kids ... but far more Chinese American and Taiwanese kids. These are just facts. I don't take it as a slam against my ethnicity that we don't do as well in math as the Chinese. There's more to life than math after all.
This is a classic case of a misplaced assumption of transferable expertise.
Mostly because of who immigrates to the US. Those Asian and Russian kids you mention probably have software engineers as parents. The Hispanic kids probably do not.
If we got the poorest and least educated immigrants from the same places, we'd be seeing rather different results.
Once you hit a certain point, sure. There will simply be people who's brains work differently and efficiently, similar to runners who have different pique physiques for their respective category (I don't think we even tap into half of that in classical public teaching but I digress). But it's a much bigger sale to say that some people simply can't pass high school level acedemics and use that against them.
The issue here seems to be that the school system is saying that the researchers aren't allowed to be a witness in any lawsuit against the school system regardless of whether it has to do with the data that was shared with the researchers.
I think a bigger issue is whether the school system should be allowed to keep any information private in the first place. If the information can safely be shared with a particular researcher then it seems like there is minimal benefit to society in letting the school system pick and choose who gets access and who doesn't.
AOL released a bunch of search queries in 2006 with they idea that they were anonymous, but it turned out you could get quite a bit of personal information from them by linking searches together.
https://www2.ed.gov/policy/gen/guid/fpco/ferpa/index.html, "The Family Educational Rights and Privacy Act (FERPA) (20 U.S.C. § 1232g; 34 CFR Part 99) is a Federal law that protects the privacy of student education records"
Now the question is if aggregated data with no PII is protected.
It’s not possible to just aggregate and be done. But it is possible to set some privacy threshold and then insure that all records conform to that acceptable risk level.
If FERPA allows sharing the data with researchers, is it right/proper/legal to share only on the condition it can't be used to harm the schools in court? Presumably that part is not in FERPA.
(And to be clear, California here didn't just say they couldn't use the data in court, they said the researchers could not testify in any court case at all against the state. But we're talking hypothetically)
The devil is in the details. These records were likely not designed to be shared and I'd assume the entire system contains vulnerabilities that could create leakage. Leakage that could be used to harm individuals in a variety of ways - from discriminating future prospects to harassment and much much more.
I agree in principle to what you're saying but we need to be truthful about what these current systems are capable of.
This is one of those things when ideology doesn't match the real world. If the amount of data is large enough and with enough parameters, removing PII doesn't do anything to protect privacy.
What about medical records? What about protected classes? What about data about vulnerable people or victims?
Student data is protected by another layer of regulation and for good reason. Also, the judiciary is a 'public' institution in general. Should we not seal records for minors? 'No exceptions' - year right..
Honestly it doesn't have to be that large. We see this all the time with data websites or apps collect. Sure, you remove John Smith's name, but you still have his GPS coordinates. For the school, you remove Professor Smith's name, but you have a professor who teaches CS 123 and has 4 graduate students. You bet you can guess who that is.
I really do support open data, especially about public institutions, but at the same time we are in an era where this information is quite powerful. Seems to make a case for something like homeomorphic encryption or something, but will that even stop these collisions?
PDF (entirely legal): https://www.cis.upenn.edu/~aaroth/Papers/privacybook.pdf
https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
But all of this is orthogonal to the core issue of whether a state government should be allowed to prevent researchers from participating in lawsuits. There is no student privacy issue involved there. Witnesses in a civil suit still aren't allowed to violate student privacy laws regardless of the data they have access to, so it makes no sense to conflate those issues.
But releases (even without patient consent, with an IRB waiver) of non-deidentified PHI data for research is allowed, and this is specifically because deidentification necessarily destroys elements that would often be necessary in research.
> You can construct some artificial scenarios where re-identification might be theoretically possible through record linkage with other data sources but in practice it's unlikely.
It is explicitly part of the HIPAA safe harbor standard that, in addition to removing the required identifiers, you cannot come up with such a scenario, and if you can, the data is not deidentified. (The last criterion of the standard is “The covered entity does not have actual knowledge that the information could be used alone or in combination with other information to identify an individual who is a subject of the information”.)
As you yourself noted upthread, you had already taken this subthread afield from that topic.
AFAICT, the thread went something like this:
The top-level concern is something like this: professors use their trusted relationship to schools in order to make bank on expert witness fees, which feels a bit corrupt and calls into question the researcher's motives.
A rebuttal to this concern is that we can side-step that issue entirely because these data sets should be public anyways (anonymized, of course!). This obviates the above concern, since the researchers won't need to compromise themselves in order to get exclusive access to data that allows them to be expert witnesses and rake in $$$$.
But the problem with that proposal is re-identification: if we can't make the data anonymous, then we all agree that it shouldn't be released (implicit in the "anonymized, of course!" caveat to "just release all the data" proposal).
Then you pointed out that even for more important data like healthcare data, FDA apparently has ways of allowing release of data that takes into account the risk of re-identification risk (I didn't know this; thanks for sharing!)
Then dragonwriter and you got deep into the weeds on HIPPA stuff.
TBH I have no idea which of you is most correct here. But anyways, there are two ways for this conversation to go:
1. You are correct, good enough anonymization is possible: Stanford researchers should not be silenced; it is problematic that they have access to data other people cannot access, but the correct solution is to negate the originally problematic distinction between those researchers and the general public by making data public. Then there is no reason for the researchers to agree to these contract clauses, because they will have access to the data.
2. dragonwriter is correct, good enough anonymization is not possible: We can go back up to the top-level concern and observe that "just release all the data with anonymization" isn't a feasible solution to this problem. Or maybe there isn't actually a problem here at all. IDK. But in any case, "obviate the problem in the top-level post by releasing anonymized data" isn't a workable solution.
Again, not following closely enough to have an opinion, but that's where we are now.
I think a good compromise position is that we should have a law stating that K12 data should be available to certain education researchers -- subject to IRB approval and so on -- without any other strings attached. Including "don't sue me" clauses in releases of public data sets does feel like an inappropriate abuse of student privacy concerns.
Whether student data de-identification is good enough or not is a total red herring. No one has accused the researchers in this case of violating privacy rules. The comments here about such privacy issues are largely hypothetical and tangential.
If you think that California needs a new law expanding research access to educational data then feel free to suggest that to your state legislators, or sponsor a ballot initiative.
It's not a red herring. It's a side conversation about a different but related topic.
Someone proposed just releasing all the data.
Someone else replied with why that wouldn't work.
Ie, a conversation happened and the topic of discussion shifted.
FWIW I agree with you on the object level question. No idea why you're being so abrasive, especially when you're the one who initiated/continued the conversational thread about deanonymization and even prefaced with "We're getting a little off topic here".
Presumably at that point you understood that the topic of conversation had shifted, and people's agreement/disagreement didn't necessarily have anything to do with the original topic... since you literally said so and no one disagreed... so your reaction here is pretty odd and off-putting.
It’s not that hard to design data release to compensate for privacy protections and statistically test for a specific level of risk. There’s a whole body of work on statistical disclosure control and there’s plenty of open source or cheap enough privacy enhancing technology available.
I’m not familiar with CA, but I expect they have someone on staff who can produce a “safe” dataset that preserve privacy and still allows for this question to be researched by low level geography, demographic, and socioeconomic factors.
That said, research that cannot be reproduced is useless. There's a balance to be struck here, and it's somewhere between "make all data public" and "lock the data in a vault".
PII is much broader than most people understand because reidentification of what amateurs would see as deidentified data is easy (often trivial), and, as a consequence, to be useful for research data is often not fully deidentified.
EDIT: As an example, the HIPAA safe harbor deidentification standard requires removing 18 kinds of identifiers, including, as one of them:
All geographic subdivisions smaller than a state, including street address, city, county, precinct, ZIP code, and their equivalent geocodes, except for the initial three digits of the ZIP code if, according to the current publicly available data from the Bureau of the Census: (1) The geographic unit formed by combining all ZIP codes with the same three initial digits contains more than 20,000 people; and (2) The initial three digits of a ZIP code for all such geographic units containing 20,000 or fewer people is changed to 000
We had a multi-month project to get a subset of our data considered 'clean', and it required a consultant, a stats PhD and many dev hours. It was healthcare, so on the high end of paranoia (justifiably) but nowhere it is as simple as dropping the "name" column
And reidentification risk can be as high as even 1% and still be acceptable for hipaa. In a dataset of a million people that’s 10,000 people identified and still be “acceptable.”
But hipaa doesn’t apply to these CA data, it’s just the clearest example of deidentification regulations I know of.
But it’s totally possible to deidentify data suitable for release to these researchers. It’s just what CA considers deidentified and if it’s still useful enough to these researchers. For the topic they are researching it should be pretty straightforward to remove PII enough to protect individuals and only remove some really unique characteristics (ie, only a single 20 year old or a particular race and ethnicity).
But I’m guessing age groups by race and gender and socioeconomic are possible to preserve without tying back to an individual. Id go so far to say as it would be non-trivial, but pretty easy, for CA to produce this for the researchers, if not to the general public.
[0] https://www.hhs.gov/hipaa/for-professionals/privacy/special-...
Maybe, though I doubt it was that easy against any but the most trivial reidentification efforts, but since most privately held PII isn’t regulated (in the US at least), there's little consequence for a social media conpany getting it wrong other than PR.
I think you could definite answer race x gender x grade but it will be harder when you factor in more unique characteristics like household income or vaccination status, etc.
In the case of education, good enough anonymization often isn't really possible. A lot of information about school students is public -- yearbooks are gold mines, as are results from extra-curricular activities such as sports, class notes about whether/where they went to college or first jobs after college, etc. Still more can be purchased. As yet more can be inferred from public data (eg home address, rough estimates at parental income, etc.). This was back in the day. I'm sure now it's much worse.
Most of the questions you want to ask about education are about treatments and outcomes. If these treatments (eg extra-curriculars) and outcomes (eg attended college, graduated HS, etc) are public then you can often figure out which student corresponds to each supposedly anonymous data-point.
Maybe not perfectly. But way more than you would think. It's like the statistics version of those little logic puzzles from grade school -- "Four people have red hair. Five are girls and two are boys. Billy and Sally are chewing gum. Boy gum chewers have brown hair. No one over 5 has red hair. Sally is 4. etc. etc. etc. Match each person's name with their hair color.". You sort of figure out a small set of data points, then look at the results of the paper and reverse engineer some statistical calculations, and then a surprising amount of the others start falling into place.
We didn't have a name for it when I did this work, but the basic point was "if you publish the dataset everyone will know little Johnny Table's test scores and GPA". Today this is called a reidentification attack.
(I don't really know enough about the topic to have an opinion on the article per se, but "just publish everything" is definitely not a workable solution :))
(I also haven't done this work in a LONG time, and there's now a whole lot of academic work on the topic that didn't exist back then, so there are probably much better consultants for legitimate organizations looking to hire for this work.)
If you want proof, you can use google to find LOTS of papers along these lines analyzing real datasets. I think https://arxiv.org/pdf/cs/0610105.pdf is a fairly typical example.
HIPAA says that 4 or more digits of a zipcode is PII. The people who protect your healthcare think most of a zipcode is too cardinal to reveal.
How many schools serve more than one zipcode? More than a few?
Exactly. With this bit being particularly outrageous:
> “Also, be aware,” wrote Cindy Kazanis, the director of CDE’s Analysis, Measurement, and Accountability Reporting Division, “that your actions have adversely impacted your working relationship with CDE, and your response to this letter is critically important to existing and future collaborations between us.”
I think a bigger issue is whether the school system should be allowed to keep any information private in the first place.
Are you genuinely suggesting that the public should have access to all attendance records, grades, test scores, etc etc of all students everywhere? That's the sort of information these researchers have.These have the potential to revolutionize private computation and analysis, as they provide provable hard (theoretical) limits on the amount of information you can learn about individuals regardless of the type of analysis performed on the proxy dataset.
I requested data about STAR, the California standardized test used for Lowell's admissions. I wanted rows of the form (randomized student ID, STAR question ID, answered correctly), however they were recorded, and literally nothing more.
They rejected the request because (1) they claimed such records didn't exist, which makes no sense because how exactly did they administer the test then; and (2) because standardized testing is carved out, in their opinion, from the related sunshine law.
Why did I want these records? I wanted to show that scoring well on tests and using them to gate admissions doesn't mean what people think it means. Specifically, that if you administered the test Lowell used (STAR) by hardest question first, then terminated the test after the student gets N (close to 1) questions wrong, you would select nearly the same list of students. Only asking the vast majority of students only e.g. 1 question, which they all get wrong, can't possibly measure how much they study, how comprehensive their knowledge is, etc. But these claims are routinely made in defense of the test and its purpose in selecting a class. This is coming from someone who wants test based admissions.
So clearly political, right? I had to carefully word my request around all these conclusions. If you read the CDE's requirements, they really have specific political goals. You either align with them or you don't. And I tried to work around that, and I stilled failed. They just looked at the absence of a political bent, and correctly concluded that it wasn't evidence of absence.
If you want to do good, politically impactful educational research: run your own school. That's what the CDE wants you to do. It's not about discovering how to improve public schools.
But at the end of the day it's fundamentally important to understand at what point transparency and privacy intersect.
> But at the end of the day it's fundamentally important to understand at what point transparency and privacy intersect.
"At the end of the day," these conversations about privacy are like 15 minutes long at private schools. People still keep sending their kids to private schools. I just don't know how much it matters.
They surely care about privacy in their internal research and metrics, but they don't employ a full time Privacist. They might employ someone who checks the right boxes for them and deals with FERPA shit. But because they are aligned with the parents in delivering the best educations, for the most part, they are trusted to do with data what they want, and that sometimes includes inviting outside collaborators to look at it, without anywhere near the same faff as the CDE.
If you're a journalist and you want to help a private school make a better education, out of the thousands of private schools, one of them will both let you write about it and also tell them something they don't wanna hear. Some might use privacy or whatever as the reason they don't want to collaborate with you, but on average it will be about trust.
The CDE is never going to do that. There's only 1 CDE, and they are there to preserve the status quo.
Very similar things happen when investigating criminal cases. There's possibly hundreds or thousands of instances of some type of misconduct or improper arrest... but none of the defense attorneys with those sorts of cases will talk about it with the press because of the very real potential harms of talking with the press. Or the ones that do talk are too high level.. or they might have some ulterior motive like self-promotion. It's really hard to express how many issues are a direct result of lawyers understandably, but systematically not raising any public awareness about truly awful things.
Have you tried getting your data through CPRA requests? I'm out in Illinois and our law is pretty decent and not super familiar with CA's public records nuance, but it's really worth a try. What I know though is that California CPRA officers get away with a strange amount of abuse of the law. But even with that, you might be surprised what records are available. So if you do submit some requests, don't exactly expect it to be easy or immediate. Expect to be stonewalled, and need to sue at some point though. But IME public record suits are pretty hands-off (except when they're not..). And most of the lawyers I've worked with are upfront about what they will and won't litigate over.
One thing you'll find is that.. basically nobody is looking into most of the awful things you'd expect would have eyes. It's very likely you'll be the only one doing those requests, or incrementally identifying how to get what you want through multiple requests over months. But each step breaks new ground and turns into feedback loops if you can build a community around it.
> Only asking the vast majority of students only e.g. 1 question, which they all get wrong, can't possibly measure how much they study, how comprehensive their knowledge is, etc.
This seems like a strawman?
Yes, a single question can't measure those things to a high degree of certainty.
But if you have students that do poorly on all the hard questions, and students that do well on all the hard questions, then asking them a single hard question might be 80% predictive of what group they're in.
Why is it bad for that percentage to be high?
The reason the test has lots of questions is specifically to increase the predictive quality. Being able to loosely predict from a small subset of questions seems reasonable to me. It doesn't mean the test is failing to measure the student's knowledge.
Aren't you clearly saying you already had a desired outcome and were just fishing for the data to confirm it? I mean, I wouldn't give you any data in that case either. It's a strong signal that you are motivated by something other than what the data shows.
What's wrong with how they want to use the data? Sort the questions, run the algorithm, see how well the scores match the real scores.
Do I think that, in principle, data sets can be anonymized? Of course I do.
Your incredulous tone and excessive ellipsis seems to imply you find this position to be ridiculous, so maybe you'd better be a little less snide and a little more expansive on, what, exactly, your problem is.
For example, if you take all students that took course A at time X, course B at time Y, course C at time Z and so on, eventually you might be able to narrow it down to a very small group, perhaps to even a single student.
Moreover, the logic would have to carry over to the very common practice of anonymizing data in professional communications (like training). Which would have HIPAA implications for some students.
The common anonymizing practices have been utilized for decades without privacy breaches of note. That track record is also what your argument would have to defeat.
Also, the burden of proof is on those that say that the data has no privacy implications, not on those who are like "ehhh, it's probably safe to release this."
That depends. The entire dataset of course because it’s everyone’s student records. But you can probably subset it to the extent that it’s still useful and perturb enough to protect individuals and be statistically equivalent.
And you could also generate a bunch of aggregate results that do stuff like identify average grade differences before and after periods while correcting for other differences without including individual identifiers.
I’m saying that the data OP can be released with perturbations made to protect privacy and the data are still useful.
I don’t think anyone is calling for the raw data to be dumped. But for as much data as possible to be released.
Yes, obviously there is a level of aggregation where privacy concerns no longer hold.
But there is no trivial transformation that allows education researchers the data they need but preserves anonymity. Education researchers want to aggregate and statistically sample the data in new ways; pre-aggregating it removes most ability to do so. If you want to do a principal component analysis of a few variables-- good luck with aggregate data.
If you provide nearly any data at the student-level, there's a pretty high chance that it can be deanonymized.
At the same time, the state's position of attempting to prevent education researchers from participating in litigation (when using only public, non-restricted data) is egregious.
For example with school attendance data you could easy release a dataset at the county level with every student’s record with unique generated student id, race, grade, gender, absences by year (or even month) and still have 5-20 of each category to be able to show attendance trends before and after Covid without being able to identify individuals. And, if necessary, suppress really unique race or gender instances (eg, maybe there’s only one trans, Native American in a school) while still being useful enough to make useful findings for general trends.
I don’t know what specific questions, but the state not releasing any data to them and claiming privacy is silly.
The census and department of ed already do this and the department of ed has a very useful description of how they apply privacy protections and validate that data are sufficiently anonymized for public release, https://studentprivacy.ed.gov/sites/default/files/resource_d...
Yes, for any kind of specific study you want to do, you can form aggregates that support it. Indeed, lots of aggregate data is already released publicly.
Lots of aggregate data is already released publicly.
If you want the actual, real data, so you can answer questions like --- "what about attendance on Mondays in students receiving subsidized lunch-- what does it predict about that student's attendance in the future?" --- you'll either need the real data, or for the state to basically do your data aggregation for your specific question.
It almost certainly can, even if it does not explicitly do so.
Unless you can guarantee privacy, which "sufficiently anonymized" does not, then no.
I'm not totally sure it's actually reasonable for a government to withhold data from researchers because they think it might be used against them in a lawsuit either. Is that a valid reason for a government institution to withhold data?
Perhaps a court case will end up establishing that the broader thing is in fact unreasonable under the first ammendment too, perhaps this is a good "test case" being even so much more egregious, you always want an especially egregious case.
I think the issue isn't being a witness in the general sense, but an expert witness which is either a paid gig or one which payment is waived because of other alignment of interests. Being an expert witness against someone you are in any kind of working relationship with is a clear and obvious conflict on interest.
> If the information can safely be shared with a particular researcher then it seems like there is minimal benefit to society in letting the school system pick and choose who gets access and who doesn't.
So HIPAA-protected data that meets the standards for research sharing should instead be made public? (And if you say, “well, its different, this is the government”—government holds lots of data protected by HIPAA.
The state is not a "someone". The state is in an extremely privileged position legally, and as such is bound by the First Amendment which you and I are not.
But it is a clear conflict of interest.
The complete ban on researchers engaging in any litigation seems over-broad, and designed to keep potential litigants from having access to anyone 'in-the-know'.
While that does seem overbroad, if the restriction were only on cases related to the data shared by the researchers, then for many cases there would need to be a demonstration that it did or didn't relate to the data, and there isn't really a way to do that without disclosing the data.
Even if you're a researcher, good quality data rarely exists. In NYC, which collects more data than any other school district, you're mostly relying on a (publicly available) 100 question survey sent to every student. The survey author must have never talked to a child because the questions are worded like a clinical psychology paper. At low income schools the survey has a 20-30% response rate.[2]
[1] https://nces.ed.gov/ccd/schoolsearch/index.asp?ID=2512750020... [2] https://tools.nycenet.edu/snapshot/2022/
I expect that they actually owe more to people actively suing them to prevent any shenanigans.
I think this is different from private instructions who have no, or very different, duty to private citizens.
Do you hold the same beliefs for publicly traded companies? Or do you just have unreasonable bars for government institutions only?
(Incidentally, I think a lot of ills in the modern world exist because of companies that exist only to increase their value in stock exchanges, rather than to be useful.)
Publicly traded corporations are very different.
My tax dollars pay for the government to operate and collect data. Not so much for publicly traded companies.
That being said, for publicly traded corporations there are regulations on what data they must release, but I think it’s mainly about financial performance.
So a private education system would not need to release anonymized data on its students. But a public education system has a legal duty.
That is not what is going on here. The research is being asked to testify against the school system by someone who is suing them.
Even if the wrongdoing is not a criminal matter, if you discover a reason that someone can be sued, then you have an obligation to inform those who could sue and act as a witness for them in court. The only exception to this is if you are the lawyer for the party you discover the data - and then you have an obligation to inform them they can be sued so here is how to fix the problem in good faith (good faith meaning if it is discovered you as a lawyer will argue that when the problem was discovered they fixed it, and thus court should dismiss the problem as an honest mistake that was corrected - the courts should in turn if not dismiss the case at least award minimal damages)
The above needs to take precedence over all contracts.
For more than a few academics, making big $$$ as an expert witness is a magnificent source of side income. (Fees of $1,000/hour, including lots of open-ended prep time, can be found.) That begs the question: Did the research lead to the desire to be an expert witness? Or did the desire to be an expert witness define the nature of the research project?
We'd need to know a lot more about the origins of this project before being able to referee this one. But if the state of California is worried about litigants using "researchers" to find and filter data that ordinarily would be available only through legal discovery processes, that's not a crazy worry.
The prep time is included in your hours. The $850 guy said he'd put in 900 hours.
(btw, it IS excruciatingly boring work. But of course, the money.)
Otherwise known as "having a job"
Apple's damages expert got paid $2 million.
One is singular.
vasco was the one responding to you.
In practice it’s equivalent to charging a higher hourly rate, but it makes billing simpler for these kinds of contracts.
900 hours is a bit more than one month. Even if he only worked 24x7 that's over 5 weeks. Assuming 10 hour days and 5 days a week that's 18 weeks, just shy of 13 weeks if 10 hours a day and 7 days a week.
As for "who can afford this" is a company worth tens of billions suing over a major product line vs a trillion dollar company.
No, I don’t have to imagine that expert witnesses do an hourly-billed contract gig but are bad billing hours.
There is, AFAIK and based on what I can Google, no universal answer.
I suppose, but do you think that was more than 9 hours? Because 9 hours is only 1% of their time.
Plus, he works for MIT. He probably needs to clear his consulting work, which could be quick or not. MIT might have wanted a percentage. And if he wanted to use a grad student to assist him in prep work, negotiating that can add up too.
There are other ways to add to the precontract numbers, but that should be enough.
And that's a huge number of hours for each round.
It's really hard to reach "a good chunk" when compared to 900 hours. I was considering saying 90 hours in my first post.
If MIT wants a cut that's a different issue, and I doubt negotiating that will take particularly long.
That guy was a star.
I also think there are many academics unwilling to serve as expert witnesses for tort lawsuits and they are different from “professional” expert witnesses.
My main recollection was that the opening statements from the lead Samsung attorney weren't that charismatic or convincing to me. I was surprised it just wasn't that... Good.
Years later I suspect the strategy of Samsung (which certainly worked, if it was the approach) was to build a good case for appeal, rather than to focus on winning the trial itself. As it turned out, apple won the trial but Samsung won the appeals.
I don't begrudge a good defense attempting to block a litigant's experts, either. However, everyone is better off for expert witnesses being motivated by fees to provide the best expert testimony. If there was something untoward about their motivation, it would be Stanford's problem.
The whole problem here is that as a soon as a researcher signs the contract, they are barred from participating in any litigation against the department even if it doesn't involve the private data they were working with. So you have a large population of experts removed from the pool, because all the experts are likely to be involved in some type of research.
> That begs the question: Did the research lead to the desire to be an expert witness? Or did the desire to be an expert witness define the nature of the research project?
I don't think these questions are productive. You can't truly know why someone does what they do. And making the suggestion that the researchers tainted their research because of the money is purely speculative and unfair.
Martin Rinard is a star. They pay him $850/hour because he testifies well, and he's done it before. He's got the credentials from MIT so juries tend to listen to him. I remember this exchange:
Apple lawyer: So that was a lot of money! Martin: It was a lot of work.
People seem to be doing some object inheritance from an ancestor post's "one month" but I didn't say that. His work would have gone over many months.
They interview him, and he writes something. Then the lawyers rewrite it. Then they all go over it, line by line. It's excruciatingly boring. I sat in on two days of the review of a different expert witness' 300-page declaration, and they had another day planned after me! They probably have a mock trial, where he practices his testimony (I'm not sure how prevalent that is).
I didn't work on Apple v. Samsung; I was just a spectator.
I don't know what an expert witness would get in this Stanford thing, but it doesn't seem to me like the spending would quite so wild.
By $NEAR_FUTURE_YEAR the only people who wind up in civil court will be victims of extortion.
In a similar vein, we need a site that lists actions taken by a state government and asks: was this in Ron DeSantis' Florida or California?
Because we're talking about part of the government.
One of the biggest things I disagree with republican on is that "government should be run like a business" no... it should not
So "running government like a business" is a way to ensure tax money is being spent effectively and efficiently
to be clear I dont think that is the best way to accomplish that goal, thus why I disagree with it, but it is not "profit" that drives that statement
It came about because far far far far far too often government programs and spending are judged by their intentions, not their actual results.
No it is not. That is reality today. Almost no government programs or spending is measured on their results.
I would love for you to prove me wrong, and show me a government program where the resolution for any failure of that program was not "we need more money"
So there was measurable learning loss from remote learning and during the pandemic.
Ok this is known in education.
The state has only relied on individual districts to make up the learning loss.
Ok so that makes sense. There is no magic bullet on fixing the learning loss issue. The state relying on individual districts taking a multi approach to learning loss .. seems reasonable.
I don’t understand the merits of the lawsuit. The state of California is already aware of learning loss and is looking at ways to address.
To be sued because the state of California didn’t do x,y,z by the paintings seems incredible short sided and unrealistic. We are still learning how to best address learning loss from 2020.
Just my two cents
I don't find it implausible that the state could have been negligent or knowingly inequitable in its learning deficit response.
A simple example would be if it suppressed internal reports about the impacts and needs, or ignored them when structuring its response.
That's a lot to protect.
By way of comparison, Florida has 69 school districts, and does measurably better across the board in providing education.
I don’t really understand why, of all the CA government institutions, the CDE finds this to be appropriate stance though. An educational office should absolutely be held to a much higher standard than this, and should at its core value openness of information and freedom of speech. The fact that this lawsuit exists at all is an indication of deeply problematic internal values within CDE that are completely misaligned with its mission and governmental function.
Florida bad though upvotes to the left.
All data in Florida from public institutions are public. There would never have been controversy in the first place. But yeah, you're right - the Sunshine laws have nothing to do with testimony.
LOL what? I can request student's grades and disciplinary records via an open records request?
What's weird is that they are being prevented from voluntary testimony on cases unrelated to the specific shared data, thus unnecessarily removing many experts from the pool.
This bumper sticker quote doesn't really track in the real world.
In your "moral" world of brutal honesty: the children of serial killers would never find work, people who were caught cheating on a test in kindergarten would never be allowed in positions of power, and people with non-mainstream interests would be sidelined in favor of those that people more closely identify with.
Is it right to hide things from people who would use that information incorrectly and to society's detriment? I think so, and that's why I believe people should have a right to privacy.
Furthermore, am I to believe these researchers are trusted not to share student PII when doing their normal academic research, but at soon as they become witnesses against the state that trust is no longer warranted? Bullshit. If protecting PII were the motivation they would not allow researchers to access that PII and publish their findings. What they're actually doing is preventing those researchers from testifying against the state. They're not protecting students, they're protecting the state's interests.
Shallow platitudes are themselves BS - populist soundbites that can be weaponized both for things you like and things you don't. There are better conversations happening elsewhere in this very comments section that don't have to lean on them as a crutch.
I don't think you'd feel the same if you were the defendant in a lawsuit, even if you had a rock solid case.
You might be completely vindicated, but bankrupted. Or, perhaps your lawyer is a dud, and fumbled the ball. Or perhaps the jury were idiots. Or perhaps the law has some unknown (to you) technicality that you end up hanging for. Or perhaps during the investigation you honestly misremember something or misspeak and the police / investigators become convinced you're guilty and spend all their time and resources trying to pin it on you. Or maybe they're just lazy, and you end up being an easy target. Don't worry, if you plead guilty you'll avoid a lengthy court battle that you can ill afford, and potential prison time if found guilty (are you that confident in your lawyer, your finances, the jury, and the legal system?). If you plead no-contest, you avoid jail, weeks or months of time off work defending yourself, and just do probation. But wait, I thought you had the Truth on your side?
-Louis Brandeis
Society shouldnt accept this data should be behind paywalls or accept high costs to access it. Or paper only releases to stop release restrictions for costs and size.
A lot of the rest I'd rather was private. Although it'd be nice to get aggregated data for certain crimes which currently are tracked at each individual department level and not in any sort of national manner.
Agreed. What gives the government the right to reject my FOIA requests for the exact specification and design files for gaseous centrifuges, implosion devices, and nerve gas?
Extreme natsec examples aside, there are a thousand reasons to keep government data private, not the least of which is constituent privacy. Deanonymizing data is far easier than preparing it for release and the data schools keep on students is particularly sensitive (I'm not claiming that that's the case with this data, just making a general observation).
Personally, I think an individual’s privacy should take precedence here.
There's no individual's privacy even at stake here. None of the data that's non-public is even material or relevant to the dispute here, beyond that the professors in question signed an agreement to access the data for unrelated matters.
In this case it was for "student-level data that detail the demographic information and the performance records over time of California’s 5.8 million students but without any names or identifying information. That data is the gold standard for accurate research. A partnership contract details the department’s commitments and researchers’ responsibilities, including strong assurances they will have security protections in place to protect students’ privacy and anonymity."
The thing about this sort of data is, removing PII from the dataset doesn't make it fully or even sufficiently anonymous. If there's only one Pacific Islander student in the Shasta Union High School District then it's easy to figure out who that is by coming it with other public data.
Quoting https://en.wikipedia.org/wiki/Differential_privacy :
] Statistical organizations have long collected information under a promise of confidentiality that the information provided will be used for statistical purposes, but that the publications will not produce information that can be traced back to a specific individual or establishment. To accomplish this goal, statistical organizations have long suppressed information in their publications. For example, in a table presenting the sales of each business in a town grouped by business category, a cell that has information from only one company might be suppressed, in order to maintain the confidentiality of that company's specific sales.
The clear justification for keeping this information private is that the government won't get sufficiently useful data without this promise. The United States Census Bureau released "confidential" information about draft evaders and Japanese-Americans; if you think they might do that again, perhaps you'll lie about some of the questions.
People who receive this sort of information are required to take special care to maintain the needed level of anonymity.
There's of course no reason why this should be used to muzzle researchers for completely unrelated fields.