Sloan Kettering’s Deal with Startup Ignites a New Uproar
nytimes.com
nytimes.com
From https://www.independent.ie/business/technology/data-sharing-...
> DeepMind, the AI lab of Google's Alphabet, has laboured for nearly two years to access medical records... Last month, the top UK privacy watchdog declared the trial violates British data-protection laws, throwing its future into question.
> Contrast that with how officials handled a project in Fuzhou... The summit involved a vast handover of data. At the press conference, city officials shared 80 exabytes worth of heart ultrasound videos, according to one company that participated. With the massive data set, some of the companies were tasked with building an AI tool that could identify heart disease, ideally at rates above medical experts. They were asked to turn it around by the autumn.
Let me make an analogy here: I advise several startups and am given shares in those companies in return for my time. A few of them make software or sell solutions that the company I work for could use. At my company I can influence whether or not we use these startups as a vendor. It is clearly unethical for me to be involved in any decision that pertains to a company where I hold an ownership stake. It would also be unethical for me to take an ownership stake after I was involved in a negotiation with those companies. I would disclose and recuse myself from any such decision. That's without adding the extra (legal) wrinkle around non-profit status, which doesn't pertain.
Disclosure is the key. As long as the people making the decision and cutting the check (investors, board, and officers) are aware of the conflict of interest, they can choose what your involvement should be.
For example, if you hire an AI expert as a consultant on evaluating AI companies, it may be with the explicit intent to buy the company for which he works as well as evaluate other companies in the space. This is one way to check people out before making an even bigger commitment.
In certain fields, there are so few "experts" that there may be no possibility of avoiding a conflict of interest because everybody is interconnected.
How reliably is the anonymization? Based on very little reading, data can at least sometime be de-anonymized. If that becomes possible 5 years from now, what happens then to the data already made public? How will individual's privacy be protected?
But, from the NYT piece, there are three issues related to the data that are at issue: 1) the dataset was generated over many decades by many pathologists who were not similarly compensated; 2) the company has an exclusive license to the data; 3) it is unclear if the patients were properly consented to allow their data to be used for commercial purposes.
The first issue is a question of money. MSK owns the data, but what about the people that generated the data? When an lab spins out a company, it's quite common for the creators of the IP get a chunk of equity (or proceeds or licensing depending on the institutional IP policy). In this case, there were dozens of pathologists that were key to building up that dataset and they were left out of the company. But this is just a money question... and could probably be solved by throwing more money at the problem. But it is important. Because, you want to encourage other doctors to contribute their findings to these types of databases to allow for future studies. Without their buy-in, these data would be lost. If you forget to include these people when that data is commercialized, you start to lose that buy-in.
The other two questions are more interesting in my opinion... Of course the company would want an exclusive license to any MSK dataset that was vital to the survival of the company. And I'm sure that MSK gets a royalty each time their data is used. But that doesn't limit the company from getting their data from another institution as well and not using the MSK data. In this case, what would MSK's recourse be? It's hard to say because MSK also owns a piece of the company... which is where the conflicts really start to raise questions. The MSK data should probably be available to anyone who wants access and has the ability to pay the licensing fee. This seems like the course most non-profits would take. But because they also own part of the company, they also want the company to have exclusive access to it. What's the proper course for a non-profit to take? What if this company could provide a truly valuable service to patients? What if the company was economically non-viable without the exclusive license? These are legitimately tough questions. I think the better solution would have a F/RAND license on the data, which may have lessened some of the above equity issues as well. However, I'm not sure the course MSK took is particularly bad. The article states that they had difficulties in getting the company funded, which underscores how difficult of a project this is.
The other side of this coin is that in most spin-out situations where the IP was created by the lab, granting the company an exclusive license to that IP (patent, etc) is extremely common. That's not entirely what happened in this case. (I don't know anything about the specifics of this case, just speculating here...) If the lab generated IP (software) that generated a model based on historical MSK data, then it's not just the lab's IP that was used, but institutional IP (generated by many other people). This makes the exclusivity a bit tougher to explain. (And is the cause of the first issue).
The biggest question to me is if the patients were properly consented for this type of data sharing. Given that the data was collected over decades, it's really difficult to know what the consent process was across the board. What they really want to avoid is the connotation that patient data is being used to fuel for-profit companies. Once that starts, you could see a slow erosion of trust of the patients... patients that have multiple choices for physicians in the NY area. Even if the patients in question were properly consented, MSK will want to avoid this type of PR to make current and future patients more comfortable.
So, in this case... the data is really the issue. Of the three issues I mentioned, I think the biggest question is that of the data access. If the license to the data wasn't exclusive, I don't know if this would have been as large of an issue.
For example, if patients knew their data would be licensed commercially, they may not have agreed to their treatment.
And even if it was legal do do this without consent, you’d still want to make sure you had patient consent just to avoid these issues at all. If they didn’t have consent, it’s a huge PR issue, even if not a legal one.
Point being, I don't see us running behind on this.
Chinese companies and government will have access to billions of medical records to train their algorithms on while their western counterparts are squabbling with individual hospitals over access to a few thousand records.
China has the talent, money and data to be the leader in this space.
https://www.healthcaredive.com/news/facebook-healthcare/5207...
https://healthitanalytics.com/news/facebook-nyu-will-use-art...
https://techcrunch.com/2018/06/15/uk-report-warns-deepmind-h...
That's a completely acceptable trade off to me. China and the west have different perspectives on how users privacy should be handled.
Under what ideology does that make sense?
For example, what if they augmented their social scoring system (a profoundly bullshit, evil enterprise) with this health data, and decide that you shouldn’t get whatever because your health data says so? What if the health data they choose to harass you with is just strongly correlated with being Uighur?
Also, it’s pretty presumptive that this data is actually valuable for the diagnostic task it’s claimed for. Biotech companies die all the time due to someone’s presumptions.
Google deepmind should just have hospital’s bid for data prices... kind of like for google fiber or amazon headquarters shopping cities.
In general, the world is better off with more information: why not open source the data if better cancer diagnosis is the goal?
Point being, hospitals have a very poor track record regarding their disposal of medical equipment and records so a motivated individual could probably easily find this data in a recycling center.
The biomedical-industrial complex in the US makes my stomach churn. So many conflicts of interest, rent-seeking, monopolies, and nepotism.