Computers do better than pathologists in predicting lung cancer type, severity
med.stanford.edu
med.stanford.edu
The study is "Predicting non-small cell lung cancer prognosis by fully automated microscopic pathology image features" by Kun-Hsing Yu et al. It's open access: http://www.nature.com/ncomms/2016/160816/ncomms12474/full/nc...
* by an un-quantified amount using only histology slides -- a limitation not enforced in actual medical practice.
(2191 so far in 2016)
Regarding the hook in the headline -- computers surpassing pathologists -- it's a bit like automated driving in that even if true the immediate problem is the social and economic system. That is, we're not going to be removing the pathologist from the diagnostic and prognostic process anytime soon for many reasons, so how instead do we leverage machine learning in concert with the human observer to improve the diagnostic system? For that reason, decision introspection may be as valuable a topic of research as improving classification accuracy: justifying a particular automated classification to the pathologist, directing them to representative regions, and describing regions of feature space in biological terms.
[1] http://ieeexplore.ieee.org/xpls/abs_all.jsp?arnumber=6945104
Given how many applications we're seeing for AI it seems that there may be a battle for attention of lawmakers and others for which AI application we put into production first given the limited bandwidth our regulators have; it's not like congress has been especially productive over the last many years.
Computer chess algorithms completely blow away human competitors. The strongest human, Magnus Carlsen has an elo rating of 2857 [0].
Stockfish, the strongest chess algorithm (open source, btw) has an elo rating of 3445 [1].
Computer chess algorithms are so much stronger than humans that if the human second guesses the algorithm -- the human is probably wrong.
You may have been thinking of this bbc article [2] in which an amateur cyborg player beat grandmaster cyborg players -- the amateurs were crunching additional metadata about what situations were best for their play. However, they didn't beat stockfish, they beat other cyborg players.
[0] https://ratings.fide.com/card.phtml?event=1503014
[1] http://www.computerchess.org.uk/ccrl/404/
[2] http://www.bbc.com/future/story/20151201-the-cyborg-chess-pl...
How is that true? I'd imagine it's to reduce the search space against human players and to make use of records of games played - a large corpus of which have involved traditional openings. It doesn't seem to imply the acknowledgement claimed?
All of this is not to say that computers have not overtaken humans in chess. They have. But primarily because they have superb qualifications as practical players - they never make crude errors and all humans on occasion do. This trumps all other considerations. But the vast gap you see when comparing the human lists and the computer lists is exaggerated - 500 Elo points means a 95% expectation. I am 100% convinced that if Magnus could be properly motivated (think big $$$ and somehow convincing him that scoring a decent proportion of draws as White would be a "win") he could deliver for humanity :-)
Interesting theory. Personally, I'm not sure it's clear that good strategy can overcome ruthless tactical precision. I'm also not sure a human could ever be motivated to achieve a 0.00% tactical blunder rate. (Much as I would love to see human strategy defeat computer tactics.)
You are 100% wrong. Computers overtook humans at chess in 1995. Unless there is some insight Magnus Carlsen or any other GM has which they are not sharing, it will remain that way.
Another factor is that it is trivial to change the computers repertoire of openings and there are a wide choice of these. Humans including Magnus require weeks to months of preparation before they are ready to play new openings or deviate from prepared lines.
Finally, (and I freely admit that this is my own personal opinion), Magnus Carlsen may not be our best choice of human to play against a computer. There is an unmistakable emotional fragility to him which manifests when he is losing (cf his games with Anand); a good deal of his strength lies mainly in the early middle game and ending but computers are superior in the latter; and he often wins games out of sheer stamina- a strategy that wont work with the silicon beast.
The surprise came at the conclusion of the event. The winner was revealed to be not a grandmaster with a state-of-the-art PC but a pair of amateur American chess players using three computers at the same time. Their skill at manipulating and “coaching” their computers to look very deeply into positions effectively counteracted the superior chess understanding of their grandmaster opponents and the greater computational power of other participants. Weak human + machine + better process was superior to a strong computer alone and, more remarkably, superior to a strong human + machine + inferior process.
It sounds like since then computers have improved enough such that humans no longer help.
[1] https://www.amazon.com/Race-Against-Machine-Accelerating-Pro...
[2] http://www.nybooks.com/articles/2010/02/11/the-chess-master-...
This is the blog post that introduced me to the idea that human + computer might be better a computer alone: https://rjlipton.wordpress.com/2015/07/28/playing-chess-with...
It's called Freestyle Chess or Advanced Chess. Humans have beaten the best computers this way, but I'm not sure it is clear that human + computer outperforms a computer alone consistently.
Incidentally, that's a great blog to browse if you like chess and math/CS theory.
http://www.nature.com/nature/journal/v533/n7601/full/nature1...
* Doctors disagreed with each other frequently. Most senior won.
* Results from one patient were given to another patient
* ... or results were lost entirely
* Lab destroyed samples before examination or stained incorrectly
* Samples never received, lost in mail or never picked up from airport
These folks did have a tough job, looking into microscopes for many hours per day. At some point I imagine things just start looking the same as fatigue sets in.
Which reminds me of Lotus Notes. 20 years ago it widely used to automate workflows. But I haven't seen anything similar widely used since.
Medicine features in it.
From: The effects of safety checklists in medicine: a systematic review: [1]
Safety checklists appear to be effective tools for improving patient safety in various clinical settings by strengthening compliance with guidelines, improving human factors, reducing the incidence of adverse events, and decreasing mortality and morbidity. None of the included studies reported negative effects on safety.
The result of such a process can be a checklist, or a workflow diagram, or a program. The form of it isn't really important, so much as the fact that the business logic being applied (by a human or a computer) to each case now embeds all the previously-scattered knowledge of the domain explicitly, making each node applying it behave as the union of its institution's knowledge-bases, rather than as the intersection of those knowledge-bases.
I'm told this is a huge problem in medicine. They might be even more hierarchical than the military.
[1] http://www.nature.com/ncomms/2016/160816/ncomms12474/full/nc...
[2] http://www.nature.com/article-assets/npg/ncomms/2016/160816/...
I presume using a pretrained net probably won't help you since the features found in pathology slides are so unlike those of natural images.
I wonder if the secret sauce that makes science effective isn't so much the scientific method, as the spirit of skepticism behind it. Without that, it's easy to make errors. And wishful thinking is pervasive because there are strong incentives: career advancement, corporate influence, ego, and so on.
My view, assuming AI doesn't kill us, is that it's going to save all of our asses (from ourselves).
Why are mammograms still being subjectively interpreted?
Are there not already companies today where you can upload medical imaging results and get back results that beat humans?
Computer-aided detection and diagnosis software is used in clinical practice, e.g. see [1]. I believe that false-postives-per-scan is still somewhat of an issue.
[1]: http://www.hologic.com/products/imaging/mammography/image-an...
Well it would be nice to first validate the results somewhere else (peer review should actually be completed) and no doubt some real world trials proceed and their results evaluated. This is what any FDA approval, at a minimum, would entail. Keep in mind what the author's themselves note in how this study may differ from the "real world":
"One limitation of this study is that cases submitted for TCGA and TMA databases might be biased in terms of having mostly images in which the morphological patterns of disease are definitive, which could be different from what pathologists encounter at their day-to-day practice."
Putting this into practice also requires looking at the bigger picture; not just the accuracy of the diagnosis vs. humans.
For example I have found that in studies doctors are much more conservative in their diagnosis than an academic algorithm with no implication on patient outcome. For instance if the question is a) "cancer" or b) "not cancer" the doctor, fearing patient harm, malpractice suits, career derailment, will be biased to identifying "a) cancer" because the perceived costs of a false positive (treating for a nonexistent cancer which may have significant ill effects) are lower than a false negative (not treating a real cancer leading potentially to death). This will reduce the human doctor's accuracy on diagnosis.
What the "right" bias is in an individual case can bring in many more factors and moral questions out of scope for this study.
This is not to argue that tools like these should not be developed and utilized but practical application is often more difficult than just solving the technology challenge. I'm certain that we'll see many advances in healthcare due to ML.
Fatigue, willpower drain, emotion, and many other factors can make a top-tier human expert into a mediocre one (or worse) in a way that is very hard to detect.
Even if computers aren't perfectly matching the ideal human (which they will inevitably pass, of course) they're still massively valuable in that they can maintain their level of expertise when humans cannot.
By which I mean, the news is not that computer > human, but that medical industry/profession/jobs are taking first baby steps at being automated away.
The test set is interesting: they used a completely unrelated dataset of tissue microarray spots. This are tiny, 1-2 mm circles. Whereas the TCGA images are full size tissue cassettes, up to 1" on the short side.