Using Deep Learning to Help Pathologists Find Tumors
research.baidu.com
research.baidu.com
It's a bit of a dirty secret in this space that pathologists have a pretty high error rate on a lot of these tasks — it's just tough work for human eyes to do literally hundreds of times every day. Applying computer vision techniques can not only improve accuracy and reproducibility over human assessment, but you can do types of analysis in seconds-to-minutes that would literally take years for a human. We're just scratching the surface.
There are lots of ML challenges here, but just as many general tech/engineering/design challenges. So if you're interested in working on bringing work like this to the masses, we'd love to talk at PathAI.
It's true that screening is a particularly interesting application because of the issues of fatigue and low true positive rates. On the other hand, decades ago (i.e. well before deep learning approaches) we had clinically approved classifiers that did better than average radiologists for some of these tasks and the uptake still hasn't been that impressive. Lot's of non-technical issues around making stuff like this standard of care.
The lack of labeled data is definitely a challenge, as you call out. But a sizable chunk of what we do is power a platform and network of pathologists to get this data within hours for training purposes. We think there will always be a very real need for human pathologists, but that the bread-and-butter work in pathology can be better handled by well-trained and thoroughly-validated algorithms.
And yeah, the non-technical issues are just as important as the technical ones:
* There's very limited use of digital imagery in clinical pathology at all. Fortunately, that's not the case in research pathology, and the success we've had in that field has been moving clinical labs toward an investment in digital pathology.
* Reimbursement (in the US) will be an issue. There are only a few options for billing payers for pathology reads, and they aren't necessarily in lock-step with the potential future of the industry.
* Like I mentioned, this opens up a class of analysis that just isn't feasible for humans to perform. It's up to us to show the value of this type of analysis.
* The regulatory environment is a real thing. We aren't hiding from this, and are creating processes that allow us to build and iterate software like we'd like to, while still faithfully meeting our regulatory burden.
So far, we've found our approach to be viable, and we've had some really strong early results with our customers (and solid revenue!). So I'm pretty optimistic, for sure.
I believe these techniques will have a huge impact on how we do pathology as well as things like screening radiology, and that part of that will be by breaking down the silos such specialties work in, at least to a agree. I also think we have quite a way to go on the technical side but it is achievable (not to do everything people dream of, but to make significant improvements).
I also think it will take much, much longer than most people on the research side believe (hope?) to even approach standard of care. These systems are not built to move fast.
I'm glad to hear you are getting good/interesting results, and hope you are focusing more on validation and breadth of data acquisition than a lot of groups do :)
In this field, this kind of technology will be augmentative for the near-term at least, so generalization matters less than elsewhere provided the false negative rate is kept low enough. Outliers can be flagged for direct inspection.
It's true that if you are assuming a human/algorithm team some things are easier, but that doesn't make the problem go away.
I trust that science, research and ongoing work will end up providing interesting results, but many many startups will burn their cash before being able to provide real world usage services.
Like you said, there are LOTs of challenges here. Certainly not a "low hanging fruit", not that a profitable business either since most countries will squash costs anywhere they can because the health cost keeps growing, and a very difficult legal environment to deal with.
However, I am very thankful for your hard work and will to push toward a better future for health.
In surgical pathology, the patient is still under surgery while a pathologist makes a rapid, preliminary diagnosis on fresh tissue. This can be done in under 10 minutes, but typically takes 20-30 in practice -- the bulk of which time is spent flash-freezing and sectioning. Pathologists have only a few minutes to review after prep, so any digital augmentation technology must run on the order of a few minutes to be of use in guiding the surgery. This guidance may be crucial because many tumors cannot be characterized ahead of time due to the lack of non-invasive diagnostic tools (the most promising is gene profiling of bloodborne cells, but that is very challenging due to low circulating concentrations). Tumor type and grade are a large factor in a surgeon’s aggressiveness-vs-risk calculation, and this is especially true in brain tumors where cognitive damage poses a quality-of-life risk which must be weighed against potential survival gains.
(not a pathologist, but I've done research, classifier, and software dev in the field -- very small, hands-on lab, so spent a huge amount of time in frozen section, OR, and with paths reviewing sections)
Just something i've been wondering about since I've got cancer on both sides of family and have been pondering doing full-body scans (which still seem quick excessive on the risk/reward)
The recent paper out from Google, "Scalable and accurate deep learning with electronic health records", has an notable result in the supplement: regularized logistic regression essentially performs just as well as Deep Nets
To truly leverage the power of machine learning, an end to end solution where the tissue is processed in a more data rich manner would be better (eg spatially aware single cell assays, non destructive thick slice imaging). This would feasibly replace the current system entirely, as it truly would do something no human could do, not just do it more accurately.
Their GitHub repo states the following: "You need to apply for data access, and once it's approved, you can download from either Google Drive, or Baidu Pan."
In this type of cancer, a lower specificity is an acceptable trade off for a very high sensitivity.
For example, a conditional random field could express "either these patches both contain a tumor or none of them does" (which is helpful when there's something suspicious on the patch boundary) and the consequences of committing to either possibility can propagate over the whole field. In contrast, a convolutional layer would have to make the decision independently for each local area.
That approach has the advantage that you'll learn about techniques roughly proportional to their current popularity, but it has the disadvantage that explanations in papers tend to be brief and you have to put them into a coherent whole yourself.
If you prefer textbooks, I heard about http://www.deeplearningbook.org/ but didn't get around to reading it. In addition to neural networks, you'll probably also want to read about classical statistics and probability theory, since that's the origin of concepts like conditional random fields, which can be mixed with neural networks but are unlikely to be covered by literature on deep learning.