Microsoft Finds Cancer Clues in Search Queries
nytimes.com
nytimes.com
The article name is: J. Paparrizos, R.W. White, E. Horvitz. Screening for Pancreatic Adenocarcinoma using Signals from Web Search Logs: Feasibility Study and Results, Journal of Oncology Practice, June 2016.[2]
[1]https://blogs.microsoft.com/next/2016/06/07/how-web-search-d...
[2] http://jop.ascopubs.org/content/early/2016/06/02/JOP.2015.01...
But it's funny, I was thinking the same thing the whole time I was reading the article... what are the search queries?!?
Second, after seeing the type of queries, I do not think that this is all that helpful. If a person has unexplained weight loss or yellow skin or eyes, they should always go see their doctor right away. My guess is that most of the specificity of this study comes from those two terms (weight loss combined with yellow skin). Just getting out that message will do a lot more to save people's lives than violating their privacy in this manner.
How do you suggest getting out the message? Blanket public health advertising would be very expensive and would be ignored by most people.
I think search histories are potentially very useful but privacy is important too. It would be interesting to see if there's a way of using search history information in a way that preserves privacy.
GPs often do not catch rare conditions or determine the underlying cause of milder symptoms.
It would sure be spiffy if mining queries could identify early symptoms, working backwards if you will, rather than be used to find cancer patients by working forward from symptoms already identified by the medical establishment.
My GP had x-rayed me, looking for a cardiac anomaly - he suspected an aortic aneurism. What he should've done is order some soft tissue radiology instead. Having your GP consider a CT or MRI is my advice.
Hope in your case it's all for naught. Best to check though.
I really hope you are better.
How is this project anything more than "other people snooping around in my search queries," or any better than simply tuning the search engine to highlight those results more if they are believed under-represented?
I think it's interesting that this research can happen.
That seems rather different than "Hey, 10 years ago you searched on some stuff that indicated you might have had pancreatic cancer. Sorry we didn't catch it sooner, since the 10 year survival rate is below 5%..."
To be clear, either one seems potentially quite cool and interesting. But if someone else besides me and Clippy are privy to these results - unless I explicitly shared them - then it seems a little creepy.
In retrospect, I suppose I wouldn't be too terribly offended if I died already, but my results were later able to help prevent someone else's untimely demise...
As it is, nearly any symptom you put in will bring up the possibility of cancer, so it’s all just noise.
Being able to connect disparate symptoms that the patients don’t connect themselves is a good thing.
Snooping on searches isn’t necessarily the greatest way of achieving that given all the privacy implications, but it may be reasonably effective.
But these symptoms map closely to other serious diseases as well, notably liver disease or liver failure. (I've had them; I know.)
While it might make sense to flag such symptoms as serious and requiring professional advice, in this case I don't see how you could possibly distinguish between the two given the vagueness of the input data.
That raises a real question about how you inform the user/potential patient without freaking the person out, or driving him or her down the wrong path.
Some other commenters have mentioned that they have been mis-diagnosed by a GP, and that perhaps this system would help. I think mis-diagnosis is a different problem. This system's real potential value is in driving more people to get professional help, who likely is a GP.
Tragic exceptions aside, most GP's are very good at what they do, which most days involves keeping "regular" people healthy. Perhaps the biggest problem in health care (at least in the US) is the lack of enough GPs in many communities and the inability of people to get access to them conveniently and affordably.
Your recent search queries suggest you may have cancer, please seek advice from your doctor.
Or some time in the future...
The pattern of your searches seems to indicate you are a terrorist, we aren't telling you and we have called the thought police.
1.They agree legally bindingly not to give any of my data to anyone.
2. They do not show me a single ad anywhere.
Obviously, that would mess the current business model of everyone counting on adwords revenue, so I am not holding my breath here.
At the moment it's the opposite.
I believe it is actually kore profitable than advertising. With ads, an average person is worth between $0.01 and $1 a month (depending on what type of ads), much less than the $9 for YouTube Red.
Results: We found that signals about patterns of queries in search logs can predict the future appearance of queries that are highly suggestive of a diagnosis of pancreatic adenocarcinoma. We showed specifically that we can identify 5% to 15% of cases, while preserving extremely low false-positive rates (0.00001 to 0.0001).
"We don't often do this, but did you make the following searches regarding the health of yourself or a loved one? ... SEARCHES FOLLOW ...
Studies show that a large proportion of people making these searches for medical purposes should talk to a doctor about these symptoms. Here's a number to call if you do not have a personal physician: (555) 555-5555"
Meanwhile, their false positive rate is as high as 1 in 10 000. Can anyone weigh in on whether that's per user in total?
If so, half their warnings are right, half are wrong. Which is not bad at all, but it is quite important for the wording of your warning.
"The data used by the researchers was anonymized, meaning
it did not carry identifying markers like a user name,
so the individuals conducting the searches could not be
contacted."
As when AOL or Yahoo released their anonymized data set, it is often easy to take someone's search history and work backwards to find out who they are. How can they ensure that personally identifiable information has been scrubbed 100% from all queries? Maybe a user searched a courier tracking number, and that info can now be looked up on the courier's site and tracked back to their home or office address. Each additional piece of info gets you one step closer to identifying who they are.Yet one more reason to use DuckDuckGo for your general search needs.
Predicting users health data based on simple searches is terrifying, I'm not sure why anyone would be happy with Google or Microsoft having this information, especially when their customers could use this against you.
Could we see a case where, when someone searches for one thing, instead of seeing results that pertain to that immediate query we see results that match common future searches?
http://www.nytimes.com/2012/02/19/magazine/shopping-habits.h...
The NYT can't even keep their site clean from virus malware laden advertisers. Ridiculous.
Google Flu is working different. They try to predict a flu epidemic by counting related search queries.
But Microsoft is predicting the health of a single person based on his search history.
Edit: Thinking about it: ofcourse Google, Facebook and others could do the same because they also gather user data.
A few years back, Google decided to recalculate YouTube view and subscriber counts to counter bot usage. They store so much information about every request that they have been able to detect views made by bots in the past, from patterns in this data.
If you collect data for one purpose, you can't use it for any other purpose, unless you explicitly and in easily readable language told the user about it before.
You cant retroactively get permission to use data for other purposes either.
And currently medical or research purposes are not listed in Microsoft's or Google's ToS.
This would have been a lot more interesting if the keywords were a lot more subtle - like a change in behavior marked by a sudden craving for salty foods or whatever.
I expect people would find this more of an invasion of privacy that the search engines doing it.
Hopefully this is a wake up call to people about how simple searches quickly build very personal profiles about you that you yourself may not even be aware of (for better of for worse).
> While five-year survival rates for pancreatic cancer are extremely low, early detection of the disease can prolong life in a very small percentage of cases. The study suggests that early screening can increase the five-year survival rate of pancreatic patients to 5 to 7 percent, from just 3 percent.
WP claims 20%[1] though a glance at the referenced source suggests that the WP summary is bogus.
So the only ones who benefit from this data mining would be health insurances who could get rid of people who'll incur very high treatment cost with low expectancy of success.
If mri's were faster, I think whole body mri's would be a decent screening tool. Problem is 1. They are expensive. 2. They take a long time to do and are uncomfortably loud 3. generate heat in the body.
There are also tumor markers for many cancers.
Screening guidelines unfortunately have to be doable to populations (lowest common denominator). More informed people with resources can do better if they take initiative (with some trade offs of time, risk).
In general, our bodies can use a lot of tuning. The more you look, the more you find. Some tuning has trade offs.
If you want to be proactive, you also have to ask your doctor for trials of particular tests or treatments. Doctors are conservative, and the first thing they will want to try is wait and see. That leaves you with may be 100 experiments you can do on yourself in a lifetime. We need to be able to do hundreds / week, to get significant progress towards making our bodies have 99.99999% up time.
The future is Star Trek type doctor, but a personal one for everyone. The major hurdles are economic and regulatory. Some physical.
According to the American Cancer Society (http://www.cancer.org/cancer/pancreaticcancer/detailedguide/...), about 53,070 people will be diagnosed with pancreatic cancer this year. The abstract says this method detects 5% to 15% of cases: that's about 2,700 to 8,000 correct detections. Assuming there are 100 million people using Bing (https://www.quantcast.com/bing.com), between 1,000 and 10,000 cases will be wrongly detected (0.00001 to 0.0001 false positive rate).
Also, the detection of 5% to 15% of cases would seem to me to refer to only Bing users; I doubt they're claiming to be able to detect 5-15% of all cases of pancreatic cancer.
Would've been nice if these things were actually spelled out in the article.
If that's true, how did they have enough real positives to measure such a low false positive rate?
Also, the detection of 5% to 15% of cases would seem to me to refer to only Bing users; I doubt they're claiming to be able to detect 5-15% of all cases of pancreatic cancer.
Yeah, I'm being dumb, there's no way that percentage is out of all cases!
Back several years ago Google Flu Trend also claimed to have 97% accuracy compared to CDC data. But later on it just found to be way off to the real data. Did the author compare their study to the Google Trend.
Also it's not clear how they achieve the conclusion of low FP. Did they randomize their sample pool and run their predictability model several round?