Paper linking frequency of search terms to violence against women retracted
retractionwatch.com
retractionwatch.com
Then the game of telephone began as bloggers simply rewrote the article and added their own spin: https://www.techmeme.com/210507/p14#a210507p14
- Ars: 96% of US users opt out of app tracking in iOS 14.5, analytics find
- iMore: 96% of iPhone users have opted out of app tracking since iOS 14.5 launched
- Wccftech: Analytics Reveal 96 Percent of Users Have Disabled iOS 14.5's App Tracking on Their iPhone
- Ubergizmo: 96% Of iOS Users Have Opted Out Of App Tracking
- …
Flurry since updated their study to make it clear what it was measuring and found opt out in rate to be about 24% for prompted users (not the 4% cited): https://www.flurry.com/blog/ios-14-5-opt-in-rate-att-restric...
But none of those articles are interested in correction and the initial visceral reaction and herumphing tweets have all since occurred. Hundreds of thousands of people got the wrong idea.
What it shows it that there is very little original reporting, people don't ask study authors for clarification, and misinterpreting of data happens all too often.
> They reported the number of hits Google displays at the top of the page as the number of searches made for that search string.
>These numbers were exceptionally high because, well, the search phrases were not enclosed in quotation marks. All this in an article arguing for the value of tech-enabled “rapid response” research. A culture of “move fast and break things” is common in Silicon Valley, but academics typically work a bit more slowly and carefully to avoid these kinds of errors.
http://searchengineland.com/why-google-cant-count-results-pr... http://homepage.ntlworld.com/jonathan.deboynepollard/FGA/goo...
In order to return fast results, there are caches and precomputed results in almost every level of a web search query.
But an accurate count implies that you will bypass all the caches and count every single document for every single query term. That's very very expensive in both CPU and memory.
We had special internal keywords that disabled caches and produced accurate results. All of them came with special warnings that querying too fast with these special options could bring down the whole search engine. A constant worry was the new employee that knew very little, learnt these keywords and then ops were dealing with crashes across ten of thousands of machines due to out of memory.
As to the fabrication. The number are not random numbers. They are estimates based on what we think is the best approximation of the real numbers.
We had done internal research to compute the best approximation function.Also, it's very difficult to figure out what is real and what is fake from the outside. I was involved in https://www.nytimes.com/2005/08/10/business/worldbusiness/ya... Yahoo was 100% correct on their claims. I run the queries myself. But external researchers could not verify the claims because all external queries were hitting caches. So they were trying unique queries in order to estimate the index size which end up having a lot of problems.
Well, that's what happens when your institution is brimming with ideologues who practice one sided research and immediately praise any results that confirms their political, dogmatic biases. Doubly so when criticizing certain results or topics will get you implicitly or ex-communicated, particularly if you are not part of an approved protected class.
The retraction doesn't matter very much, the damage has already been done, and far more eyes will have been exposed to the results than to the retraction.