AI tool cuts unexpected deaths in hospital by 26%, Canadian study finds
cbc.ca
cbc.ca
In this case, the absolute risk when measuring for death in the GIM pre-intervention and GIM post-intervention are 0.0215 (2.15%) and 0.0146 (1.46%) with an absolute risk reduction of 0.0069 (.69%).
While the relative risk is 26% across the pre- and post-intervention, the absolute risk reduction is only 0.69% with a NNT (number needed to treat) of 1/156. Which means that 1 patient in 156 was helped by this intervention.
In addition, they had 2 false alarms for each true alarm and could suggest that interventions were performed in patients who did not require it — more tests, medications and possibly increased risk from said interventions.
This shows that the CHARTwatch ML/AI is not helping at all that much clinically.
Many researchers go to relative risk because it shows better results!
It came to roughly the same conclusion as the gp comment when provided with the study PDF.
https://chatgpt.com/share/66eb09e3-7a74-8008-afa8-3b60161d24...
(Though obviously this approach still requires you to go and look at the PDF yourself to make sure it isn't making anything up)
Could is doing a lot of work in letting you interpret what it's saying however you like.
Happy to talk about it.
Are you in the Healthcare industry?
It was bs, as usual.
That is always the issue with alarms. You have a fine line to walk. Too many alarms and people become complacent and learn to ignore alarms. Too few alarms and you don't draw the attention that is needed.
That's enough to anchor "alert == I might find a problem" in the user's mind.
Instead, the failsafe has the effect of merely invalidating the current test, and making the appliance unable to run a test correctly until either power cycled or the appliance's developer executes a secret series of commands that are not shared with us.
So of course an operator of the appliance found a way to feed in a false "I'm here!" with a loop, to trick the appliance into never going into failsafe…
That's for ~6.8% of all tests being false-positive, ~93.2% being true-negative, and ~3 tests that should have triggered failsafe did not.
I don't if you meant it as a counterpoint for what I said, but it really isn't.
The problem is that given any positive at all, the chance it points to a problem is still virtually zero.
If it was 6.8% of all tests as false positives and 2% true positives, probably people wouldn't have silenced the alarm.
If it goes off 8 times a day and 2 of them are true positives, then people have recent memories of having to fix problems pointed by the alarm.
The reduction I am arguing against is: "Historically, extra information and diagnostics that have an error margin results in worse outcomes because we misapply it; therefore don't build these systems."
And the nurses who want decent pay and can do travel nurse, do travel nurse
https://www.sciencedirect.com/science/article/pii/S277265332...
the headline says we're talking about death: does that mean 1 life was saved for every 156 patients?
>In addition, they had 2 false alarms for each true alarm and ... and possibly increased risk from said interventions
but wouldn't this study have captured any deaths from those interventions, so the 1 out of 156 life-savings was net?
this study was measuring deaths and what you are suggesting would be outside this study, but it could be measured also.
I found this interesting:
> 1 truly alerted patient for every 2 falsely alerted patients was deemed an acceptable number of false alarms
In an ideal world, the nurse to patient ratio would be high enough that patients could be seen on regular rotation frequently. I've never been in a hospital where this was the case. So a system that can correctly prioritize resources for critical cases even if it's pulling resources away from non-critical cases will probably result in a net improved outcome.
If your home alarm caught one real burglar per each three occassions it triggered, I bet you wouldn't develop alarm fatigue. I certainly wouldn't.
https://en.m.wikipedia.org/wiki/Multivariate_adaptive_regres...
I think there are two causes of Wikipedia maths articles' general awfulness:
1. They're probably written by people that just learnt about them and want to show off their superior knowledge rather than explain the concept.
2. The people writing them think it's supposed to be a precise mathematical definition of the concept, rather than an easy to understand introduction. It's like they're writing a formal model instead of a tutorial.
Often the Mathworld articles are a lot better than Wikipedia, when they exist at least.
The term "MARS" is trademarked and licensed to Salford Systems. In order to avoid trademark infringements, many open-source implementations of MARS are called "Earth".I think this is the part that people miss the most. When a purchasing decision is made based on something like "who has the best quality shoes in price range X", competition can occur.
When the buying decision is "will I live or die", there's not really any choice made there. Couple that with the complete lack of transparency for how much a give procedure will cost, and you've strayed so far away from a free market that it's not even recognizable.
I mean, even the hospital can't even remotely accurately tell you how much something will cost before you actually get a bill...
Folks think that removing the "profit motive" will somehow cripple the whole system, or hospitals will try to save any penny they can (spoiler alert: they already do)
This is kind of the reason the Japanese economy is stagnant and continues to fail in winning global marketshare. Businesses that are too good will fail or at least not compete with businesses that settle for being good enough.
...where "good enough" is relative to a particular level of quality and price point, of course. Otherwise there wouldn't be different markets for rich and poor people. And this mechanic helps avoid a "collapse into mediocrity" that you'd otherwise get if all goods and services were offered at a single price point.
The real problem is what you identified at the end, that healthcare isn't anything like a free market. There's no buyer mobility, no transparency as to the level of the service you're getting - heck, you don't even know how much you're going to pay in advance, unlike almost every other industry.
Increasing prior authorizations, increasing paperwork complexity, increasing hold times on the phone, obfuscation for who is responsible for what, constantly changing coverage so people have to change providers, and otherwise discourage them from seeking care.
The Baumol effect you link to only shows that wage demands from health care workers go up in proportion to the wages of other workers. This means (roughly speaking), that reducing the health care budget will reduce the effectiveness of your health care system, because you're able to afford fewer people (I think this is the point you're making, please correct me if I'm wrong).
But that's entirely the point of starving the beast! By cutting funding to some federal department, that department becomes less effective, which makes people think that the government is incapable of running said department, and makes them open to the idea of privatizing the department. Et voila, you've opened up a whole new market that can be exploited for profits! The holy grail is opening up a market with inelastic demand such as health care, where people, no matter what you charge, will be forced to buy your product. This program has been incredibly successful in the US, which can be seen by comparing their health care system to that of other wealthy nations.
Even declaring that is the case doesn’t change that it’s still clearly a personal judgement depending on the individual.
"Depending on the individual" here means, depending if you're a share holder, or the patient dying on the cot.
If you can get 20% by paying... what, presumably <5% more for some ML tool that double-checks stuff and flags risky stuff... perhaps it's something we want to do.
It's entirely possible that we want better healthcare outcomes - all the historical trends point to that - but that we're more or less out of ideas how to get there on the cheap. This might be a new possibility.
In your model, why do we get improved, costlier insulin if the old thing was good enough? Because we actually want to pay more if it works better, and it doesn't mean we cut something else to make up for it. You just pay more in taxes in a subsidized model, or pay more at the pharmacy with private healthcare. There's a drug manufacturer profit motive in there, but it holds true in the added-cost ML scenario too.
So, to answer your question about 'why do we get improved, costlier insulin if the old thing was good enough' it is because the healthcare system will make more money on it. If they take a % then they are incentivized to use a more expensive version and they can justify it with the word 'better' even if the person is actually worse off as their financial situation deteriorates and they and their families are forced to cut quality of life everywhere else. They put their line for good enough at the point that makes the most value for them, not the point that is best for the patient.
Outcome is not one thing. The patient wants better health. The provider has an interest in profits. The government has an interest in optics…well anyone using “AI” does.
That being said, I'm fine with a reduction of resources if additional resources don't increase the quality of my care. In Canada, doctors don't really like to prescribe antibiotics for minor infections.
Americans find this bizarre, but for a minor infection antibiotics are going to screw up your stomach bacteria and long-term health to maybe treat a disease that your body can easily handle on its own.
There's no magic value that comes from allocating resources to a problem. Oftentimes spending money has zero or negative impact beyond virtue-signalling that you care about the problem.
I don't think we should ever take any sort of superior position on this. The same motivations and outcomes occur.
Having said that, efficiency is good, especially with an aging population that will require more and more care. Resources are limited, so applying them in the most effective, efficient way possible is always a win.
Our system has major problems, but we spend less money and have a healthier population. That definitionally means we're more efficient.
> The same motivations and outcomes occur.
Our hospitals don't have shareholders that capture excess revenue as profit. Efficiency gains in a non-profit hospital typically get reinvested into the mission of providing healthcare. Efficiency gains in a for-profit hospital often go to the owners.
"Efficiency" is also measured differently in a non-profit context. A business measures monetary return on investment. A non-profit organization measures the monetary cost of achieving its mission.
Many for-profit hospitals in the United States offer free mental health clinics. These clinics have been accused of baiting patients into saying something suicidal as a tactic to involuntarily commit said patients.[2] Because appeals of an emergency mental health order are difficult, this is an extremely efficient way of making money (the hospital gets to bill the patient for their stay).
I don't believe this could happen in Canada. The goal is to get people out of the hospital because there aren't enough beds.
[1] https://en.wikipedia.org/wiki/Comparison_of_the_healthcare_s...
[2] https://www.buzzfeednews.com/article/rosalindadams/intake
As to the mental health holds, here in Canada we have a problem with social workers encouraging difficult cases to consider medically assisted suicides, which is pretty disgusting. We have people dying on waiting lists. We have people having to go to the US to get basic imagining of probable cancer cases.
Universal healthcare is superior -- again assuming proper funding, which jurisdictions like Ontario are far, far short of -- but in the current state of the Canadian system, I would never imagine bragging about it online.
Aren’t most hospitals in the US technically non-profit, though?
This is exactly why a structure like the UK NHS which is going for "what's the most healthcare I can get for the country with a fixed pot of money" is a better setup.
For instance, in the UK the female contraceptive pill is free to whoever wants it. Because that is a whole lot cheaper than extra (unwanted) pregnancies. Similarly the NHS has spent money on reducing smoking because that's cheaper than dealing with the health effects.
Smokers also help keep pension/social security costs down since they pay into it but don't collect out of it or do for much shorter period.
From a nation which should know better after being so very thoroughly roasted by Mr. Swift some few years ago: https://www.gutenberg.org/files/1080/1080-h/1080-h.htm
Abundant contraception encourages and promotes promiscuity
> the NHS has spent money on reducing smoking because that's cheaper than dealing with the health effects.
Reducing tobacco usage makes more room for nicotine OTC and vaping to replace it. Among other stimulants.
I don't think data supports your claim that tobacco use was merely redirected to other forms of nicotine. But even if it did, that's a success since they're less harmful.
And? Nicotine itself is not particularly dangerous and might be even neuroprotective if consumed in moderation. Vaping as a consumption method might be problematic of course, but I don’t think there is any research showing it to be even as remotely as harmful?
> white blood cell count was "really, really high"
You don't need AI for this.
I wish they would provide a more compelling example.
If a study found that letting cats roam hospital hallways reduced unexpected deaths by 26%, I think that would be reported, too.
"A difference-in-differences comparison between GIM and subspecialty units demonstrated no statistically significant difference in outcomes"
I've been that programmer more times than I can count. I'm much happier about being able to work on better problems instead than I am worried about AI taking away my rice bowl.
These are life-long software engineers, just like others reading this comment, using the best tools at their disposal to engineer lifesaving software. They're not using "regex" to develop algorithms for monitoring patients (???), and frankly that suggestion is so wild that one has to assume you don't know anything about algorithm design at all.
An LLM literally hallucinates incorrect answers by design and struggles to get extremely basic math and spelling correct.
You're welcome to put your literal life in the hands of a hallucinating english generator, but when it comes to healthcare, I want a "0% LLM" policy. LLM's will be the cheap things that offer substandard care to poor people, while the wealthy and elite enjoy personalized and human-centered care.
LLM's and accuracy in one sentence in the context of quantifying thresholds is stunning.
LLM's don't have a concept of numerical accuracy.
And you don’t need Dropbox for file sync. Machine learning makes integrating automation easier.
Forgoing a decade of income to get some letters beside your name selects for people who don't take orders from Clippy unless you market it well.
I don’t like calling everything AI, but I’m even more irritated by people that don’t understand the value of simple ML models for low hanging fruit decisions like the one shown here
A 26% reduction in unexpected deaths, apparently.
Using AI to find patterns in patients and intervene was something I worked on in my last job in Specialty Pharma. Theres many red flags on patients long before they even start treatment, sadly income is one of the largest red flags here in the States.
We were able to perform interventions earlier and improve outcomes with a simple regression model that tried to determined the number of missed doses.
If a loved one is in the hospital, stay with them as long as the hospital will allow you to.
Medical isnt science, and its frightening.
The weirdest thing I've experienced as a patient is that Physicians will urge you against second opinions or having multiple doctors.
Hope telemedicine becomes more mainstream, I'd like to avoid US physicians as much as possible.
I don't think we discourage second opinions, except maybe in some for-profit structures. The bad idea is to have multiple people making decisions in parallel. I'm not in the US, though.
Regarding advocacy, I don't think it's so crazy. It's very good to have a valid interlocutor when the patient is diminished. Also, hospitals are big systems with limited personalization. If someone's there to call out the system when it's trying to shoehorn too hard, it's also very good.
I would have had no problem intellectually getting through the program but quit after the first night in a hospital.
Anyone sitting at a desk can not understand how tough and miserable a nursing job is. Everyone is basically miserable and stressed out. The work is completely thankless, disgusting and dangerous with personal liability on the line if you make a mistake. Everything that we take for granted in an office setting just doesn't apply in a medical setting.
I eventually just went back to a bullshit project management job, for more money than a nurse of course. This is obviously part of the problem.
It is easy to complain about the system when it is someone else who has to help grandma to the bathroom. There is no easy solution for any of this given the demographics. It is basically a disaster.
Medical professionals, mostly nurses, are spread extremely thin. They are so busy and/or jaded that they often neglect to show any compassion or empathy until they see somebody else doing it. Having a family member nearby also keeps them accountable.
I have seen it personally too many times.
Marketing will abuse any term they get their hands on, and certainly "AI" has been abused, but in the field it usually the umbrella term for all areas of research into making "intelligent" behaviour. Be it expert systems, logic systems, machine learning, statistical machine learning, or otherwise.
If you want the details, call it a regression model. If not, why insisting on communicating the details?
Maybe there is some nuance for things like a patient in for liver issues where their liver enzymes are expected to be abnormal, but identifying when it is abnormal for them.
I’m not sure how an alarm for “high white cell count” should have had so much impact. Here in China once the doctor prescribes a finger blood test, we sample finger blood after lining up for 15 minutes, and the result is available within 30 minutes. The patient prints the results from a kiosk and any patient who cares enough about their own health will see the exceptionally high white cell count and request an urgent appointment with the doctor for diagnosis right away. Even in normal cases we usually have the doctor see the report within two hours. Why wait several hours?
> While the nursing team usually checked blood work around noon, the technology flagged incoming results several hours beforehand.
> But in health care, he stressed, these tools have immense potential to combat the staff shortages plaguing Canada's health-care system by supplementing traditional bedside care.
This sounds like the deaths prevented by this tech are caused by delays and staff shortage and what this tech does is to prioritize patients with serious issues? While I appreciate using new tools to cut deaths, it looks like the elephant in the room is staff shortage?
Like I'm not sure what this measure means, it's not like 26% of people that would die in the hospital would be made immortal or something.
This stuff will not happen because its good technology that can save lives. Rather, the public pressure from AI performing better at saving lives that humans.
The anecdotes of 'oh it was wrong that one time', will pale in comparison to success. Maybe Insurance companies will be the winners and be our advocate. I've already seen medical professionals use 'that one time it was wrong' as a way to ignore technology.
I am also reminded of Dilbert's PHB decreeing that all future unplanned outages must be announced at least 48 hours in advance.
So the blood was collected and labs done but it wasn’t scheduled to be reviewed until later?
Seems like a win-win. For those saying you don’t need AI, the alternative would be either across-the-board thresholds for flags for each line item (too many false positives) or manually setting it for each patient (too intensive).
The article would have been stronger with those numbers. But I wouldn’t be convinced that a high WBC count for an average ER visitor would have been sensitive enough to trigger an alarm. The prior knowledge that it’s a cat bite is important.
that's just an alert, not ai
Nobody really expects AI to save terminal cancer patients or 90-y.o. cardiacs. Unexpected deaths, on the other hand, are really nasty, both for the next of kin and the doctors themselves. If an apparently viable patient suddenly drops dead, everyone asks what went wrong.
Reducing such deaths by one fourth is a good job.
But in this instant, it's machine learning in the form of regression analysis: Multivariate adaptive regression spline
So they're understaffed and could just look into the results more, oh wow what an use of computational power.
I suppose there is a risk they will downsize more. But this is like thinking cameras were bad because they reduced the number of security guards needed to secure an area. No?
That is, I'm willing to chalk up use of "AI" as a descriptor being an editorial choice. Agreed that it isn't impressive just because it is AI, but it does still seem to be a good use of computational power.
I took your tone to be a bit of push back on this being a good use of compute.
People wonder why folks hate doctors or get “white coat” syndrome. Same shit from dentists wondering why everyone hates them.
I'm not sure where exactly you evaluated this based on (personal experience I suppose?) but this hasn't been true for me in Spain with either public healthcare nor private. Don't remember it being like that in Sweden (public healthcare) either, and I'm sure there are plenty of other European countries where the waiting time isn't significant either, and you also get great care.
Some countries seems to just have figured out how to make healthcare costs manageable, with great care, well educated doctors/nurses and also relatively low waiting times. I'd probably still say they're underpaid, because they're literally saving people's lives, but I guess that's true for everywhere, even the US.
https://www.dr-bill.ca/blog/career-advice/doctor-salary-us-v...
Sure, there's the exchange rate, but it's still quite good. The disparity for tech workers is much greater.
That and the schools simply won't graduate enough of them. Doctor shortage is a serious problem. But so is nurse shortage post-COVID.
System here seems to be in crisis. Combination of many factors.
But all my experiences in the last few years have been... very positive? Excellent recent care for my teen at McMaster Children's Hospital. Family doctor 5 minute drive away, can get appointments quite quickly. So, yeah, it's regional and situation dependent.
Otherwise, you're hooked up to monitoring equipment...
Another way of thinking about it is that non-AI systems must always perform a task correctly, or we'd say that they have a bug. Conversely, an AI system performs tasks in situations where there is some measure of uncertainty or subjectivity, and they might arrive at a way of performing the task that is suboptimal, or even entirely inappropriate, without being buggy - for these systems we'd say that they did their best given the circumstances.
In the case of this hospital study, if they had used a simple "beep if measure goes above X" system, that wouldn't have been AI, but they used an ML model which integrates many interdependent factors over time [0] and while it has a significant ratio of false positive triggers (and as such is often wrong), it applies what would absolutely count as "reasoning" in trained human nurses.
[0] "The deterioration prediction model was a time-aware multivariate adaptive regression spline (MARS) model (Appendix, Sections 1–4). The model is made time-aware by incorporating risk score predictions from earlier in the encounter, the change in risk score since the previous assessment, and summaries of changes in the risk score over time." https://www.cmaj.ca/content/196/30/E1027
How is this not a good use of compute?