It's not just us. The NIPS machine learning for healthcare workshop had hundreds of attendees this year from both industry and academia: https://www.nipsml4hc.ws
If you're an ML researcher or engineer and want to use machine learning to save lives, feel free to email me. I'm brandon@cardiogr.am. Happy to talk about our company or point you to relevant research.
I do work on data from persistent AF patients. Specifically, trying to predict AF recurrence after treatment with electrical cardioversion. Basically, electrical cardioversion is an effective treatment for some subset of persistent patients, but for another subset it is not. Doctors have a hard time deciding which patients can benefit from electrical cardioversion and which will not. If we can build a model that predicts this, we can avoid unnecessary procedures (which always carry some risk) and explore other treatment options instead. If this works well, it would directly benefit individuals.
But the data isn't actually collected -- e.g., I was told Holter monitor recordings are not kept, only looked at and discarded if nothing is found; and data which is collected is often useless for automated analysis.
As someone who has worked in ML, I'm said that's the case. As someone who also worked in security, I dread the day this data is properly collected. It will not be properly anonymized, it will be available to shady people for the right price, and it's more likely your enemies and insurers will know when your heart is going to fail before you do.
NHS IT is fundamentally crippled.
18 months ago I submitted a proposal for some software I had ready to go, happy to discuss costs but suggested a small amount (£10/user/month IIRC) with a costed business plan showing the savings it could make. They declined it, fine, but then gave the proposal document to an in house developer who's been working full time on it since and still hasn't even shown a line of code to anyone.
I'd tried to get national level innovation funding for that so the local organisation wouldn't have to pay for it. NHS innovation money is only available for proven software. By which I mean having a full system tested by clinicians and ready to go isn't enough; it has to already be in production use to be innovation!
The last couple of months a marketing agency have been trying to sell some software I've written across the NHS to help people collaborate, but they just can't do it. They've concluded it has to be sold to one small area at a time, literally starting in a GP surgery for something that makes most sense rolled out across the organisation. Before I approached the marketing agency I'd contacted the front door email of about 5 NHS IT organisations claiming to help suppliers improve the NHS about how to start the ball rolling; none of them got back to me.
There's a clue to the fractured nature of the NHS in the article linked elsewhere in the comments about Google getting AI data; they've only been able to get the data for 1.6 million patients. The NHS deal with that many patients in 2-3 days.
A month or two ago their email hit national headlines for going into meltdown after someone spammed most of the staff with the mailing list in CC and they all ended up replying to all to ask to be removed from the list.
I could drone on. But maybe you get the point.
Fast forward to today, and there is more openness. I did skim the paper mentioned here, but did not see any links to the actual data, which is a shame.
My understanding is that 70% of the difficulty medical research is collecting the data. So once you've done that you want to make sure you've protected your investment of time and energy.
From an individual stand point I understand why they do it. From a societal standpoint it's such a waste.
I'm guessing they don't, which makes findings, kinda, suspect.
Also their trove of data is not a one shot. They might collect data a mountain of data on Leukemia but publish one study on Leukemia's correlation with high power lines. Publishing this paper wouldn't necessitate opening up all of their data from either a horizontal or vertical perspective.
I see your point about using the data for other research. Hopefully reviewers of the paper, at least, get to look at the data.
Of course the major stumbling block of access to good data remains.
0) https://www.newscientist.com/article/2086454-revealed-google... 1) https://www.nytimes.com/2017/01/09/technology/medicaids-data...
(edit: formatting)
A previous company I worked for was scared to death of holding any records that may be classified as health records because of the regulatory implications.
The big-data regime isn't the only regime. Probabilistic models can encode doctors' prior knowledge and also be trained with only a few dozen data points.
A famous(in medical circles) cardiologist Eugene Braunwald once said after a stint practicing in mexico: "...they have the patients we have the technology..."
The US is not the world.
What sorts of data sets are you looking for? It is probably available from other countries or could be more easily collected. I have an interest in this area too
It's there: https://en.wikipedia.org/wiki/Department_of_Defense_Serum_Re...
Or at least the samples are, and they are all comprehensively linked to healthcare records as well.
Would be expensive to analyze though, and the political implications would be huge.