The Parable of Google Flu: Traps in Big Data Analysis [pdf]
gking.harvard.edu
gking.harvard.edu
However, these 1,152 data points are highly temporally correlated (the value at time t is a strong predictor of the value at time t+1). As such, they aren't "worth" very much when building a model. Therefore, the effectively independent number of points (what you generally need to build a good model) is in reality much smaller.
Looking at the graphs, it appears that for each flu season there is a pretty regular spike of flu incidences. This spike could probably be summarized effectively by four parameters:
- Mean
- Standard Deviation
- Skewness
- Kurtosis
Looking at it this way, the number of "independent" data points of a flu season is not 365 (one for every day the year), 52 (one for every week), or even 12 (one for every month). Instead it is only 4.
Thus the total amount of data they had to build their model was 4*(# years). When you have such small data sets, you can't expect to obtain good results.
TL;DR. Time series are very hard to use in predictive models. It is often impossible to generate good predictive models based off them.
A combination of GFT search behavior and actual CDC numbers produced the best model for the authors of this paper. The (media) spin is "GFT as a stand-alone flu tracker fails".
Perhaps the article would be of a different tone if the researchers had full access to Google's data. I don't think the mentions of irony, failure, misleading and hubris are deserved nor offer a fair portrayal.
I just attended a talk that graphed the rate of physicians' diagnoses versus the truth, and the conclusion was that about half of physicians drastically overestimate, too (and it's far more than double, closer to five times more!).
So when 80-90% of people visiting the doctor for flu don't really have it and those doctors often make type II errors, you can hardly expect their diagnoses to be a reliable source of information.
I guess the flu epidemic is just a conspiration created to distract us from the fact that nobody walked on the moon or something.
On a more serious note, what's exactly the deal with misdiagnosis? If the real thing is another virus, it could also be a virus that we should treat in a similar manner? Or the doctors are loosing precious time often with a wrong first diagnosis in a significant number of cases?
Beyond that, no, it's not likely to be a virus treated in a similar manner. If by manner you mean "drugs". Oseltamivir (aka Tamiflu) only works against influenza, so if you have RSV or some other virus, it won't do you much good. The other way to handle viral diseases is to let them run their course and manage the complications, at which point a strictly accurate diagnosis isn't particularly necessary.
TL;DR: It's very hard, fairly expensive, and the payoff per-patient isn't strong.
Gist of it is, Google Trends overestimated flu cases by about double for 100 out of 108 weeks, according to the paper.
Edit: There is actually an interesting podcast available on it without a paywall: http://podcasts.aaas.org/science_podcast/SciencePodcast_1403...
Well, there's a simple solution to that... Just halve all of the results from the current algorithm!
The question has always been "Can we match CDC estimates faster and with different information?"
The emerging answer is: "No".
That's a pointless and irrelevant comparison.
There's a cost to false positives, e.g. creating more vaccinations than necessary; fortunately I don't think the CDC or any other government organizations made any policy decisions based on Google Flu. As a person who works with a lot of data, I think applications like Google Flu have a lot of potential to do a lot of good. But it sounds like Google wasn't properly rigorous in their experiments if Google Flu is such a failure.
CDC tracks hospitalisation based on lab reports, GFT tracks search terms
The assumption, which admittedly Google is making, is that these two are correlated by the same amount each year. However, in periods of economic downturn without free health care, people are going to search more and seek actual medial assistance less. Given the global economy has been shot for the last six years....
> CDC tracks hospitalisation based on lab reports, GFT tracks search terms
This is the key difference, and after 5 years it is clear that this is not a valid way to predict the flu, or even monitor it. I think it was a great experiment, and thought it was pretty creative when announced, and now it should either be retired or replaced with a new hypothesis. Because as the Mythbusters would say, this myth is busted.I think, actually, the core idea is a terrific one.
From [1]. [1] also has professor and doctor state that her models improved after adding GFT data. Data from search behavior does not have to replace CDC data. It can improve it. No myth.
[1] http://www.npr.org/blogs/health/2014/03/13/289802934/googles...
I have never gone to the doctor when I had a flu. If it had lasted more than one or two days, or if I had a temperature higher than 102, then I would have gone. But as it was, nobody ever knew that I had the flu except myself and my family.
Even if I had wanted to go to the doctor for every time I got a flu or a cold, it takes at least a day or two to schedule a doctor's appointment. By the time the appointment had rolled around I would probably have cancelled, since the flu doesn't usually last very long.
This being the US, there are also a lot of uninsured people who can't go to the doctor even if they wanted to. For those people, Google is the only option. The CDC can collect as much data as it wants from the laboratories, but you can't collect what isn't there.
When you add in the uninsured people and the people who get better before visiting the doctor, it's not at all surprising that the lab estimates are 2x lower than reality. In fact, I would kind of expect them to be even lower than that.
Error in the Google model is by definition the difference from the CDC levels.
[edit: For example, the author of the paper could have compared Google and CDC versus a third data source, like sales of cold medicine. That would have been an interesting comparison, and might have shed some light on which estimates were a better reflection of reality.]
Things like "sales of cold medicine" and the like fall under the heading of "syndromic surveillance", and believe me, they've been tried for the better part of a decade without yielding much that's impressive. Which is frustrating, because theoretically they should be amazing, but they just aren't, especially when asked to predict rather than back fit.
The CDC numbers (which are part of ILINet) are pretty decent, and one of the things they do, which hasn't been mentioned here, is they're not trying to find all cases of influenza. They have a network of physicians in all 50 states reporting "Influenza-like Illness". These people then also send samples in for laboratory confirmation of virus type, drug resistance, etc. The numbers are also backwards revised as more data comes in, meaning they improve over the course of the year.
It's a very good system for getting an idea of how "bad" the flu season is - the main problem is that it's slow. The ideal, and what people were hoping from Google Flu Trends (and other systems like it, it's not alone) is "ILINet, but in real time".
It also fails fairly spectacularly on a local level when compared to health department data from some top flight state agencies, like NY.
There are three datasets here: the CDC's, Google's, and reality. Some folks here seem unable to differentiate between the CDC data and reality. But in fact, I really did have the flu, even though CDC didn't think I did. Reality is what matters.
Maybe there is reason to believe that CDC's data is closer to reality than Google's. If so, let's hear it.
The CDC is unapologetic about their data being estimates - but they're very solid estimates, and match more intensive but smaller scale studies pretty well. But we don't need to capture every single flu case - no surveillance system will ever do that, nor need to.
Edit: I see how the CDC's numbers could be a lower bound, but not upper.
Then perhaps you didn't have strictly the flu, but some other viral upper respiratory tract infection instead.
It is hard to judge from symptoms alone (which is why lab confirmations are required for diagnosis), but as a general rule of thumb influenza A infections in adults tend to be much more severe and long lasting than what you describe [1], with distinctively higher malaise, fever and incapacitation than your regular winter cold. It is so debilitating that you probably would have gone to the doctor in any case, I would guess.
Is it possible the CDC estimates are off by half? Two main points come to mind:
* How many Type II errors do doctors commit?
* How many people with the flu do not visit a doctor? People with good health coverage are probably more likely to visit the doctor preemptively when they are feeling unwell, potentially skewing the results.
Also, while google may be off in the exact measurement when compared to the CDC, it looks like the shape of the graph correlates with the CDC data. Overall, this looks approach promising.
/newest can be so random.