FastMRI initiative releases neuroimaging data set
ai.facebook.com
ai.facebook.com
> The application process includes acceptance of the Data Sharing Agreement (found below) and submission of an online application form. The application must include the investigator’s institutional affiliation and the proposed uses of the data. NYU fastMRI data may be used for internal research or educational purposes only as described in the data use agreement and may not be redistributed in any way without prior permission.
> Read and agree to the data use agreement below to apply for access.
Seriously?!?!?!?!?! You call that open source?!??!?! OK, let's leave the semantics aside for a moment. Yes, a dataset like that would be very interesting and I would happily play around with it for a few weeks and see what I can come up with. If it's something worth exploring further, I'll happily document it and open source it. But that isn't something I can really estimate without being to explore the data and see first hand what it is. For that reason alone, myself (and dozens off the top of my head) will roll their eyes and pretend like it doesn't exist.
False claims, over-hyping with no real understanding and bureaucratic crap like this is what is slowing everyone and everything down, the sooner people understand it, the better.
This is not an isolated case - large portion of those "open source" data sets are "Please sign up, send us an email and mail a DNA sample". And I'm talking about data sets which hold 0.0 personal or classified information.
Spare-time tinkerers aren't really the intended audience here:
>>The neuro data set will allow researchers to test their models with data from additional machine types, new sequence types, and different coil configurations that were not present in the previously released fastMRI knee data set. Radiologists also look for different diagnostic properties (such as contrast in texture between different neural tissue) in brain MRIs. These differences present an interesting and challenging machine learning problem to solve and will help researchers develop models that generalize to more clinical settings.
It's unlikely in the extreme that anyone in that audience will be stopped by a simple data sharing agreement. And it's also unlikely in the extreme that anyone outside that audience will know what to do with a bunch of raw k-space MR datasets. Domain knowledge is an absolute necessity with this data.
Same applies to DNN's(if not more so) - take any large DNN with the papers, data set and code to the author and ask them to give an explanation as to why it works as well as it does while it performs terribly on a different data set, even a similar one. "Well yeah, it's curve fitting which works here but doesn't work there". Why? ¯\_(ツ)_/¯
My rant concerns a different problem - if you have such data and you want to share it, just go ahead and do. 1 of every 10000 might do something useful with it but we aren't talking about nuclear experiments where something can blow up, are we? Worst case scenario someone's cpu or gpu might overheat, big deal. Just ditch the entire bureaucracy crap, we have enough of that as it is in our daily lives.
Nah, it's the other way around. Specifically in this domain, innovation is definitely driven by domain experts (researchers, medical equipment manufacturers), and they contract out or hire technical expertise as they need it.
>To put it as question: do I really need to explain how open source works in theory and practice on a place like hackernews?
Your definition of open source is much too narrow and does not include what these researchers meant by it.
>Especially given that most of the world's infrastructure in practice runs thanks to the collaboration of millions who have built most of what we use in their "spare time" as you call it.
Certainly not in medicine.
>Same applies to DNN's(if not more so) - take any large DNN with the papers, data set and code to the author and ask them to give an explanation as to why it works as well as it does while it performs terribly on a different data set, even a similar one. "Well yeah, it's curve fitting which works here but doesn't work there". Why? ¯\_(ツ)_/¯
You've identified the precise reason why all this "machine learning" stuff has been such a dud when it comes to medicine. That is not good enough and never will be, particularly the "works so well on the dataset you overfitted on while performing so terribly on other datasets".
>My rant concerns a different problem - if you have such data and you want to share it, just go ahead and do. 1 of every 10000 might do something useful with it but we aren't talking about nuclear experiments where something can blow up, are we? Worst case scenario someone's cpu or gpu might overheat, big deal. Just ditch the entire bureaucracy crap, we have enough of that as it is in our daily lives.
They spell out the reasons for requiring a data sharing agreement in the data sharing agreement itself. They want you to cite it, they don't want you to sell it, they want you to confirm that you understand you're getting a dataset with no warranties, FDA approvals, etc.
And as I said before, the chances of you doing something useful with it may be 1/10000, but the chances of you doing something useful with it while not being motivated enough to sign a simple agreement and wait for the download link is essentially 0.
I think any application of deep convolutional neural networks should be alongside a radiologist. If we speed up scans and make up for it with convnets it is very hard (practically speaking: impossible) to properly validate that they will not hallucinate away rare abnormalities. It will also be impossible for radiologists your spot errors like this in the wild because of the reduction in quality of the scan.
What happens when the scanners change their behavior in some subtle way that is unaccounted for by FastMRI? It could start erasing a ton of subtle abnormalities and this would not be possible to check for since the original scan will be lower quality.
There is just too much redundancy in MRI data, and initiatives such as FastMRI are fundamental for us to learn what the limits are of feasible acceleration. Also, some MRI scans take forever and cannot be used in vulnerable populations because of, e.g., breath holds, the need to stand still, etc. The image quality, perhaps counter-intuitively, in some situations improves with acceleration.
It’s interesting research for sure. I hope it stays far away from actual clinical use for a while, for the reasons I highlighted. I’d like to see convnets work alongside radiologists for a while and prove robustness to dcanner changes in the wild before we start shoving them deep in the stack where radiologists can’t review what’s happening.
Such that i am not sure the risks are the same as say convolution nets reconstructing large brain structures.
K-space data is also saved for reconstructions and processing later on, though everyone prefers to avoid that as it’s horrible and lots of storage is required.
I’ve also worked at a university site where the raw data was collected and used on a daily basis, but that is presumably less common.
Rather than dreaming about that, the focus should definitely be on "clinical decision support", i.e. "something useful that will save a radiologist some time and won't just get in the way". Not too many examples of that exist right now. Even speech-to-text is not a solved problem in their domain.
Radiologists are notoriously conservative for that very reason. Dreamy-eyed image processing and computer vision researchers have been trying to get radiologists to abandon some of their caution for many decades. All in vain. "Hard pass", they say most of the time. The dream usually doesn't get as far as their malpractice insurers, but it certainly dies there if it makes it that far.
The benchmark for this technology is not perfection. The benchmark is human radiologists. Yes, this technology will miss things, so do humans. But if it's performing better than the humans we should prefer it, even if it's not perfect.
I think you have misunderstood the application, possibly because of the way the parent framed it.
This is not a project to interpret MRI data, it is a project to apply ML to accelerated scanning, i.e. inferring data that is not actually measured.
So it's a real problem, if a systematic bias attenuates some signals that would be interesting, there will be nothing there for a radiologist (or other ML system) to perform on.
Think of this as more of a "algorithmic super-resolution" approach.
I agree the validation and regulatory path is really problematic for less common presentations.
It's probably more interesting to use for artifact detection / QC issues, especially when you do it quickly enough to initiate a re-scan. More problematically but still interesting to do artifact removal. Acceleration is tricky though, for the reasons you mention.
First, this work is about taking data in the sensor domain ("k-space") and reconstructing it into an image. Doing this with partial k-space data and hand-coded heuristics is a completely standard part of the MRI research agenda and has been for quite some time. See, for example, http://mriquestions.com/k-space-trajectories.html. Further, several of these techniques have already made it into routine clinical work, and this acquisition-side stuff generally happens before the radiologist even sees the image (reliable acquisition is in the interaction of radiographer with the scanner manufacturer's software).
There's also various claims here that seem to imply learned reconstruction inherently implies the risk of hallucinations without recourse. Naturally, one should be careful about this, but it's just a matter of careful cross validation: hold out examples of abnormal anatomy for the test set. There's other ways to attack this problem too: training can be done partly or mostly on synthetic data because we have reasonably good forward models of the physics. In this case, one could choose a wide variety of arbitrary synthetic anatomies during training, to further ally the fear of always hallucinating the "typical human brain" from any scan.
Slow acquisition and image artifacts in MRI are a fact of life for people in the field and I believe there's huge scope for improvement if we had more intelligent reconstruction and acquisition. Ideally the reconstruction would feed dynamically back into the acquisition to gather more context as necessary; the MR machine is, after all, one giant programmable physics experiment. This is already done in a limited way, but in what I've seen it relies on a lot of hand-coded heuristics. And guess what's the logical step after hand-coded heuristics? Yes, learned models where you objectively optimize for a final result, rather than hand-coding based on a few examples.
Final note - publicly releasing human data is a massive effort in data cleaning and careful anonymization. Not to mention that the acquisition of each sample is extraordinarily expensive. So bravo to these guys for going to the effort.
It's a natural research direction to accelerate imaging; MRI is fundamentally limited in its data acquisition rate. You are essentially giving protons a wack into an excited state and extracting spatial information by listening to how they decay. There is a time constant associated with this decay that just comes from the physics and can't be reduced, and you have to wait for the decay to be more or less complete before you can give another wack and acquire another slice.
MRI already relies on classical compressed sensing to overcome this limitation and make imaging feasible at all, but there is good reason to think that you can do better than classical compressed sensing by making use of prior knowledge of the physics of the imaging process and the structure of natural tissue.
People with spine issues, MS, Lyme, and Lupus for example usually refer to it as "limbo land". Neurological diseases are notoriously difficult to diagnose and often go without an official diagnoses for months, years, and even decades.
Thanks!
EDIT: As for viewing, they come with a proprietary viewer loaded on the CD. I do have the ability to export them as JPG image stacks though.
I've been working in medical imaging and deep learning for years, and have recently become disillusioned by the technology's potential to disrupt the radiology industry. But I wonder if there aren't alternative use cases for e.g. educational purposes. I'd love to know more about what you wish you could do or know! Please feel free to email me.
It's a serious technical challenge but the benefits could be enormous.
IIRC the vast majority of the cost of an MRI is the amortized cost of the imager, so faster scans should hopefully directly reduce the cost to patients, perhaps to the point that regular full-body MRI scans for preventative healthcare could be feasible.
This is an interesting problem. Speaking with a number of physicians among the family, there's a perspective that having a population performing all these tests can possibly cause more problems than they solve. If the goal is to holistically make a person healthy and happy, discovering diagnoses that have no practical effects on someone's well-being can result in reducing their well-being just by knowing about it. Humans are notoriously bad this way. Tell someone their liver is somewhat different from an average adult liver and they'll start assigning symptoms to it.
I'm more thinking, wow you could really do a lot in the way of automated diagnoses if you had longitudinal data sets like that.
This is certainly false for clinical MR imaging in Canada. The biggest component is the cost of interpreting the scan, i.e. the radiologist's time.
"regular full-body MRI scans for preventative healthcare could be feasible."
What's the point? What are you looking for? What is the false positive rate on that? Screening for the sake of screening is a mirage[0].
[0] https://sciencebasedmedicine.org/a-skeptical-look-at-screeni...
https://ai.facebook.com/blog/deepfovea-using-deep-learning-f...
Take 10s to read yourself instead of complaining about others not reading for you. It's at the top of the page currently, under a heading, "Apply for Access".