AI Detects Heart Failure from One Heartbeat: Study
forbes.com
forbes.com
This paper offers like 1/1000th the evidence of that one.
It's got a lot of interesting applications for our field which I am excited about, but there seems to be a tendency among non-experts to consider it a magic bullet that can solve any sort of problem. In particular, I am concerned about applications where conventional approaches have already converged on an optimal solution that's used operationally, but somebody wants to throw AI at it because they thought it might be cool without first understanding the implications.
From the abstract:
This study’s findings suggest that skin markings significantly interfered with the CNN’s correct diagnosis of nevi by increasing the melanoma probability scores and consequently the false-positive rate. A predominance of skin markings in melanoma training images may have induced the CNN’s association of markings with a melanoma diagnosis. Accordingly, these findings suggest that skin markings should be avoided in dermoscopic images intended for analysis by a CNN.
There was a hn discussion about it a while back: https://news.ycombinator.com/item?id=18415031
> Aircraft landing
> Evolved algorithm for landing aircraft exploited overflow errors in the physics simulator by creating large forces that were estimated to be zero, resulting in a perfect score
> In an artificial life simulation where survival required energy but giving birth had no energy cost, one species evolved a sedentary lifestyle that consisted mostly of mating in order to produce new children which could be eaten (or used as mates to produce more edible children).
> (Tetris) Agent pauses the game indefinitely to avoid losing
>Agent kills itself at the end of level 1 to avoid losing in level
>Robot hand pretending to grasp an object by moving between the camera and the object
"Finasteride is a compound that is used in two drugs. Proscar is used for prostate enlargement. It is old, out-of-patent, and has cheap generics. Propecia is used for hair loss. It is a newer, and (at the time) very expensive. The only difference is that Propecia is a lower-dose formulation.
What people did was to ask their doctors to perscribe generic Proscar, and then break the pills up to take for hair loss. Doctors would then justify the prescription by "diagnosing" enlarged prostate. This would enter the patient's health records. If you apply deep learning without being aware of this "trick", you would learn that a lot of young men have enlarged prostates, and that Proscar is an effective, well-tolerated treatment for it.
Health records are often political-economic documents rather than medical."
They didn't show that.
First of all, why just one heartbeat? You never capture just one heartbeat on an ECG anyway, and "Is the next heartbeat identical to the first one?" is such an important source of information, it seems completely irrational to exclude it. At least pick TWO heartbeats. If you're gonna pick one random heartbeat, how do you know you didn't pick an extra systole on accident? (Extra systoles look different, and often less healthy, than "normal" heart beats, as they originate from different regions of the heart.)
Secondly, why heart failure and not a heart attack? One definition of heart failure is "the heart is unable to pump sufficiently to maintain blood flow to meet the body's needs," which can be caused by all sorts of factors, many of them external to the actual function of the heart - do we even know for sure that there are ANY ECG changes definitely tied to heart failure? Why not instead try to detect heart attacks, which cause well-defined and well-researched known ECG changes?
(I realize AIs that claim to be able to detect heart attacks already exist. None of the ones I've personally worked with have ever been usable. The false positive rate is ridiculously high. I suppose maybe some research hospital somewhere has a working one?)
(For comparison, the "CHF beat" looks a lot more like a healthy heartbeat.)
I think it's a sort of academic machismo. "Look what we can do - isn't it amazing?"
I saw the same thing in Robotics recently. An academic came to give a talk on localisation using computer vision: they cross-referenced shop signs that were seen by a robotic camera with the shop's location on a map to get a rough estimate of where the robot was. My first question was "what is the incremental benefit of this approach when it was combined with GPS?". It turned out that the researchers just hadn't used GPS at all - almost like they considered it to be "cheating".
I feel like many academic disciplines have unwritten 'rules' that you need to follow if you want to be included in the conversation. Not all of those rules are sensible.
Ninja Edit: N=~30 patients. For something like ECGs which are readily available, they really should have tried to get more patients. A single clinic anywhere does than 30 EKGs per day. Suggesting this is clinically applicable is ridiculous. It's way too easy to overfit. Chopping up a time series from one patient into 1000 pieces doesn't give you 1000x the patients.
I even think this approach probably will work. Very reasonable given recent work from Geisinger and Mayo. But why are ML people doing press releases about such underwhelming studies?
Well you're on the right site for it!
At least you didn't claim that you could come up with something better by tinkering on a rainy Sunday afternoon.
I could agree with your claim if you meant bootcamp programs or data science sorts of coursework, but machine learning is generally grounded in both measure theoretic probability theory and a robust understanding of applied statistics before moving on. After that will be the basics of pattern classification, clustering, regression and dimensionality reduction. Last of all will be very domain-specific tools for NLP, computer vision, audio processing involving e.g. deep neural networks.
Clinical research that isn’t making ridiculous claims tends to get much less press.
Furthermore, of all places to lap that crap up... this hacker news site is frankly one of the worst.
I mean look at this submission. Yes it’s true people are pillorying it here (including some doctors), but i don’t recall much interesting well designed medical research being discussed here (though arguably maybe not the place for it)
Its not that difficult.
Basically, they took 18 electrocardiographic tracings (sampled at 128 Hz) from participants without CHF, of whom 13 come from women. They compared them to 15 electrocardiographic tracings (sampled at 250 Hz) from participants with CHF, of whom 4 come from women.
Hard to even know where to begin with this one.
The study avoids the obvious pitfall, which is to put different slices of one patient's data into both training and test. The press also reports the training accuracy (100%) when the test accuracy/sensitivity/precision metrics are all at around 98%.
Another encouraging sign is that when you dig into the 2% error rate, a majority of those errors turned out to be mislabeled data.
The study also acknowledges the following:
"Our study must also be seen in light of its limitations... First the CHF subjects used in this study suffer from severe CHF only...could yield less accurate results for milder CHF."
I think this is a good proof of concept but that the severe CHF and tiny sample size (33 patients) means that we're a long ways away from clinical usage.
There is nothing to see here.
The case dataset is independent of the control dataset.
The missing experiment is to have a third dataset from yet another machine, with both positive/negative examples, and use it as the test dataset. Then transferability questions are at least somewhat addressed.
https://www.reddit.com/r/MachineLearning/comments/dj5psh/n_n...
The main impact might be that if this holds up people could be tested with a short hook up in an office instead of with a 24-hour monitoring where they have to bring back a Holter device the next day. Of course, that 24 hour dataset may have independent value of its own for further diagnostics beyond just whether the patient has CHF.
The datasets for positive cases and negative cases come from different databases. n=30 patients, on top of it.
All this does is recognize the patient/ECG technician who recorded the data. It's basically certain it doesnt generalize
In a past job I did a combination of manual and machine-learning-based analysis of cardiac signals. We didn't have ECG, but did have PPG (blood flow) and PCG (sound) signals, and a pretty large study group. I recall there being one study participant who's signals were very clearly indicative of heart failure, enough that we raised the issue with our medical advisor about whether the subject should be deanonymized and contacted. In the paper they state that "the CHF subjects used in this study suffer from severe CHF only"; my suspicion is that a simpler, "hand rolled" model based on the features of the ECG could compete very well with this CNN approach for finding the same level of pathology in the ECG signal, without the "black box" of a CNN casting doubt on the technique.
Along with doing a lot of good and making a lot of early catches, I suspect that relying on AI to do medical analysis is going to bring into sharp relief just how much medical science DOESN'T know about the human body and its mysteries. I think we're a long, long way away from handling medical science over to AI and the real fun of AI-guided exploration is about to begin.
If someone here wants to start a project with AI on top of ultrasounds, I'm all in.
let me know at hn at angra.ltd and I can give more details
We, however, have access to the report too.
Also how can the detection of a progressive disease be 100% accurate? I guess details ruin the click bait too.
I found a great startup opportunity for you.
- Recruiter