Everyone gets numbers wrong, even the New York Times
climateer.substack.com
climateer.substack.com
This is the core rule I learned as a researcher. More often than not, reputable journalists will get it more right than wrong on aggregate. On average, the numbers published by reputable journalists will be right and not misleading.
These are some interesting examples, however, the core rule: "One source is not evidence" is probably more important here to remember. Read a variety of sources and articles. I often find that sometimes in the same paper they will contradict numbers. Then I know, the editor missed it.
The other possibility is a chain where manipulation can take place in a non obvious way. For example reviews of PC hardware like CPUs on release day. It might appear as if you are reading many different sites with differing review methods but in practice all of them were likely provided a CPU and motherboard by the same manufacturer chain. Both AMD and Intel have been caught in the past providing golden samples for reviews and motherboard manufacturers auto overclocking boards. SSD manufacturers providing MLC for review and then switching to TLC NAND once it hits retail. There are so many scandals that involve this sort of thing that its very hard to trust reviews that aren't retail sourced and recently too.
Reading multiple sources of news isn't enough, you have to analyse where it comes from and how it could have been manipulated and falsified based on where the chain of custody of the news item comes from, a much harder problem and largely impossible given how the news hides its sources.
I haven't traced it but the New York Times and Google seem to have similar problems.
Usually it's better to look at the state's website or the CDC, but they have different dashboards that often don't calculate the same useful metrics for you.
Or at least, they'd be useful if they were right.
If you are ever very close to the subject of a story, to the point where you know the important details, and then you see that subject get covered by the New York Times, you will see just how bad they are. Exceedingly, horrendously, tragically bad with facts, and trying to put an ideological spin on everything. But I'll grant that it's often less bad than some of the others.
https://www.johndcook.com/blog/2021/01/18/gell-mann-amnesia/
Despite being called out that the data was being misrepresented, they never updated the article.
The idea that Russia and a Ukraine supply 25% of the wheat in the world and they don’t correct that is egregious.
The tweet's claim is "...the statistic is misrepresented. The 13% number only captures people who worked from home BECAUSE OF CORONAVIRUS." That's exactly what it says in the story.
> Do you have examples of stories
And the weasel words don’t help here. Once the damage is done, a retraction cannot undo it.
Wen Ho Lee was in solitary confinement for nine months under threat of possible execution for things he did not do based in large part on drama whipped up (and never fully retracted imho, not that that matters one fucking bit for the nine months and other impacts on him and his family) by the early NYT reporting on his case.
The one in the example may be a brain fart, but so many other errors could be prevented by standardizing a few things ... decimals and thousand sepatators is another one
For example, in French, 1 million dollars is "1 milion de dolars", but 1 billion dollars is "1 miliard de dolars". Then, 1 trilion dollars is "1 billion de dollars", and a quadrillion dollars is "1 billiard de dollars".
https://en.wikipedia.org/wiki/Long_and_short_scale (there's a map of current usage, most of Europe uses milliard etc.)
I developed a rule of thumb over the years: if you can't Fermi estimate some metric or quantity within an order of magnitude, you probably don't understand the metric very well. If you can't estimate it within two order of magnitudes, you probably don't understand it at all.
I believe their mistakes are equally likely to be intentional as they are to be simple mistakes. Especially when it comes to topics where there is an agenda at play.
But even factually correct numbers, presented neutrally, can still fuel an agenda. For example, most corona related articles focus on the number of deaths, but don't mention the number of cases nor the amount of people with long-term effects, nor the long term effects of e.g. the organ scarring that the 'rona can cause. It's deception - or downplaying - via omission.
>> I believe their mistakes are equally likely to be intentional as they are to be simple mistakes. Especially when it comes to topics where there is an agenda at play.
> To use the case in the example, what is the agenda for under-reporting, by an excruciatingly obvious 3 orders of magnitude, the number of Covid infections?
It's the NYT's agenda to downplay COVID, obviously. /s
IMHO, misinterpreting mistakes as intentional lying is a common tactic to justify ignoring evidence and sources that contradict preconceived beliefs.
It's true that pretty much every kind of media is biased in some way, but I think a lot of people let the perfect be the enemy of the good and overreact to that fact by disengaging with certain reliable sources or with the media entirely. The ironic thing is that often results in relying on even more on biased and unreliable stuff, not less.
This is exacerbated by short-form social media applying higher point values to witty snipes at a piece over in-depth engagement.
That's a huge claim to make without any evidence at all, interesting opinion I guess.
1. Traditionally, editors write the headlines and sub heads. The editor knows even less about the subject than the writer so what are the chances they will get it right especially when their goal is not accuracy but attracting readers.
2. Twitter, the water cooler for all journalists, has put on public display and has quantified journalist relevance, popularity, and influence. Blogs did this to a certain extent but Twitter really put the gas to the pedal. Some news orgs force their journalists to use social media. This means that there are very real personal and career incentives to make the story fit the in-group narrative or be lambasted for it. This means that finding an angle on a story that highlights a specific narrative is the goal. To be clear the narrative is usually some worse case scenario that scares people and gets them to click. It doesn't have to be some sort of script handed down by a conspiracy.
3. Most people only read headlines and maybe the first few graffs. Often clarifying or information countering the headline is mentioned at the end of a piece. I often will read the end of an article first before subjecting myself to the manipulation found at the beginning of articles.
4. When success is about having clicks and shares there is a _strong_ incentive to publish stories before anyone else which is antithetical to carefully ensuring accuracy, integrity, and reason.
All in all, journalism is simply organized hearsay and I never automatically believe anything I am being told. That is not say it is useless, but it is important to understand its nature before consuming it.
Because newspapers still often have to fit for space and are less likely to write multiple headlines for digital and print, there are constraints that often lead to rapid iterations and are far more likely to be out of sync with the article itself.
There are times when people are communicating approximations and feelings, where imperfection is fine, and times when they're communicating precise, technical, numeric concepts where an imperfect model is insufficient.
One interesting application of neural networks is the creation word embeddings. ML models are trained to place words in a vector space, which is useful for measuring distance between words, performing arithmetic on words, or finding the closet word. Using an embedding allows your to formalize the "distance" between words, and perform fun tricks like King + Woman = Queen.
Citation needed. Is this based on any sort of formal linguistic/anthropological reasoning or just like, your opinion?
Also, the article's whole point is "people are bad at numbers" and your argument that "linguistically, all numbers close" isn't really true. Why don't rhyming words in English feel "close"? Also, "duck" and "fuck" are both verbs involving a thing you would do with your body, so why isn't that closeness?
If you're going to point out that language feels arbitrary by making arbitrary points, you're going to be "right" but you're not actually saying much. Language can be both specific and arbitrary (it's a means of expressing both objective and subjective concepts) so an argument in favor of doing your best when seeking to be objective seems pretty reasonable.
This result is intuitively obvious to me, which this article illustrates. Even if the resulting sentence is not factually true, a true-sounding sentence can be constructed by taking a sentence with the word million and replacing it with the word billion (in a majority of cases). This isn't true with duck and fuck, and duck is a noun and a verb used in generally different contexts than the word fuck.
> “Half a million known virus cases”, huh? During the Omicron peak, the US alone exceeded that many cases every day. Clearly they meant half a billion.
I wonder if the New York times knew exactly what they were writing. Both of course are technically correct. And billions/trillions/etc get into unfamiliar/meaningless territory for some with a million being the largest meaningful number so could have more emotional impact as a headline.
That may be so. But, the New York Times can't use that as a excuse for making simple mistakes, nor do I think it wants to. Language is primarily what it relies on to represent ideas. Accurate description of the world through language is the core of its value proposition to subscribers.
Edit: I originally had one extra letter, thanks jrd79
Maybe bbillion pronounced b-billion would be better.
And to add confusion, note that UK English had
1e6 million 1e9 thousand million 1e12 billion (million million)
I think this has mostly consolidated to the US usage now
That is true, but not enough to change the numbers and certainly not enough to make up for the uncounted cases. But when these media outlets get the numbers wrong, it bolsters these ideological positions built around “they’re lying to you”.
> That is true, but not enough to change the numbers
By definition, if the counting methodology is wrong (in the misleading sense), then they necessarily DO “change the numbers”
> and certainly not enough to make up for the uncounted cases
The uncounted cases aren’t deaths. Therefore it’s not a question of “making up for” those cases.
In fact, if anything, that makes it worse.
Using a singular value “total number of deaths” is generally misleading when trying to justify public health policy on large populations. A global population of 7-8 billion means that even huge numbers can be insignificant, eg 1 million over 8 billion is only 0.0125%. (Disclaimer: Numbers chose at random; this example does not represent any actual statistic around COVID).
If the cases are being overestimated then it drives the overall impact down.
Numbers and methodology both matter intensely when we’re talking about using the numbers to justify public policy changes.
The thinking goes that if someone is just a little sick they aren't going through the hassle of getting an official test and being counted. If someone is nearing deaths door they are going to take action and go to the hospital or clinic so you're likely to have a very high percentage of cases leading to death being reported. Add on to that the fact that people who died in a car accident and happened to have a non-symptomatic case of covid was added to the death numbers. That might not be a huge number but it will have an effect at some level. I don't see how it could not unless you completely ignore them.
Inaccuracies in data can cut in all directions. In this case though, over counting deaths and under counting cases (a deliberate choice made by the CDC) made the virus seem more deadly than it likely was at the time.
Is there a plausible argument for an under count of deaths? If so, I have not had the pleasure of hearing it.
One of the fringe benefits of being a programmer is that if you try, you can start developing an intuition for how things across a huge span of orders of magnitude add together. That chart alone covers 8-and-a-bit orders of magnitude differences for simple operations, and then we may want to perform those operations thousands or billions or quadrillions of times.
I don't do super-high-performance systems, but I still encounter coworkers in our normal routines who are off in the mental models by multiple orders of magnitude w.r.t. to how much something should cost. I see overestimation more often than underestimation.
(My opinion on that is that if you "think" a database query should take 50ms, when it actually does you don't worry about it and dig into why. When the answer is, you're missing an index and doing a table scan for something that ought to be 5 microseconds, a full 4 orders of magnitude difference, it's easy to not notice, because in absolute, human terms, 50ms is still pretty fast. Make a dozen or two of those mistakes, even in otherwise very large systems running lots of code with a lot more than a dozen things going on, and it's easy for overestimates to become self-fulfilling prophecies and to accidentally build systems bleeding out orders (plural!) of magnitude performance without realizing it. Underestimates are much more likely to slap you in the face and get resolved one way or another.)
Another example: I'm a fan of time-tested wisdom of all sorts, but sometimes time does move on and invalidate things. "It all adds up" used to be true when everything we humans dealt with was in the same rough orders of magnitude, but it's not always true anymore. It doesn't always all add up. If I've got a 500ms process, do you have any idea how many nanosecond things it takes to even bump that by 1%, let alone add up to anything significant? If your "1ns"-range code has effectively no loops, or is O(n) on some small chunk of data or something, it's inconsequential. There's plenty of other places in the modern world, which spans more orders of magnitude than the world used to, where this time-tested wisdom can be false, and it in fact does not "all add up".
This is one of the reason you must always profile your code if you want to improve performance. People have always not always been perfect at finding bottlenecks even when our machines didn't casually span 12 orders of magnitude, but any developer no matter how experienced can be tempted to blame the complicated code dealing in microseconds but miss the simple-looking code hiding milliseconds.
This covers mostly the first bit of the article, but if you practice this sort of sense can start helping with a lot of the other innumeracy issues encountered too. We have good practice grounds for this in our discipline.
I agree with your point. I’ve spent way too much time responding to code reviews where the reviewer asks me to be more efficient in something that takes triple digit nanoseconds per request on internal systems that typically get dozens of requests per day.
I would love a version of this classic xkcd that covers events that take nanoseconds and milliseconds [0]. I suppose it would have to also account for the difference in value between a human’s time and a machine’s unless you assume there is a human waiting in real time for every process to finish.
Same with any reporting about a police report that doesn't include a ... copy of the police report.
Funny if it were!
I mean the "at least half a million" figure is not factually incorrect, it's just... misinformation? Downplaying, either intentionally or accidentally? Same as the "probably more" line, it's downplaying a fact by adding a bit of insecurity. It's an ass covering opening paragraph.
Trader Joe's, for instance, makes signs for their produce with crazy prices. Five bananas for a penny! Yes, you read that right! They either don't understand decimal places, or they don't realize that it's not legal to advertise incorrect prices.
Their signs clearly say .19¢. I pointed it out to them, and they looked at me like I'm crazy.
Perhaps they'll care when I insist they sell them to me for that price, then complain to the county's Office of Consumer Affairs if they don't.
I don't know about American law, but here in the UK you would lose that complaint. A price label is what is known as an "invite to treat", and a store is under no obligation to sell the item to you at that price. If they say it's a mistake they can and you have no right to say otherwise.
A price of 1/5 ct would probably be considered too low to be sensible and would not be considered binding.
For example, Wisconsin:
98.08 Price refunds; price information. (1) A person who uses an electronic scanner to record the price of a commodity or thing and who sells the commodity or thing at a price higher than the posted or advertised price of that commodity or thing at least shall refund to a person who purchases the commodity or thing the difference between the posted or advertised price of the commodity or thing and the price charged at the time of sale.(2) A person who sells a commodity or thing and who uses an electronic scanner to record the price of that commodity or thing shall display, in a conspicuous manner, a sign stating the requirements of sub. (1).
https://glitchndealz.com/glitch-laws-by-state-pricing-error-...
Did you actually do either of these things?
Ignorance of math is a bad thing!
> The store must honor the price in the advertisement, even if it is wrong, until they correct the misrepresentation using the same advertising medium and/or by corrective signs in the store.
> Items sold in a grocery store must ring up at the lowest displayed price. Food stores that are in the waiver program must offer consumers one of the items free if it scans higher than the lowest advertised price.
[1] https://www.mass.gov/guides/a-massachusetts-consumer-guide-t...
Does anybody reading the sign not understand what it's trying to communicate?
"It's a banana Michael. What could it cost, 10 dollars?!"
You say "obvious" but I don't think that means what you think that means. Why is it obvious? Why would it be impossible that a store sells bananas at a loss in order to pull people in (what people in sales call a "loss leader")? If you don't buy bananas regularly, why couldn't you assume this is just a great deal?
If there were a sign advertising 5 bananas for a penny (or whatever) with no asterisk, fine print, or mention that "terms may apply", I would be livid if I couldn't take advantage of that price.
If you would like links: Here's a 5 year old piece about Amazon giving away bananas: https://www.foodandwine.com/news/amazons-free-bananas-disrup...
Here's a 2 year old coupon for free bananas from some couponing blog: https://www.thecouponingcouple.com/free-bananas/
Another blog, more free banana coupons: https://spoonuniversity.com/how-to/10-food-budgeting-tips-i-...
TL;dr- Don't put up bullshit pricing signs and assume people will understand the "joke", you're coming from a place of different cultural awareness and it will burn you when people like me get involved.
Also, grocery stores sell a lot more than just groceries.
More of the backstory: https://verizonmath.blogspot.com/
Honestly, I think the correct policy is probably just to interpret .19¢ as a syntax error.
Trader Joes could sell for .19 cents and require a minimum purchase of ten pounds, for example.
That method of representing price has been normalized.
We all know what it means.
That's different than miscommunication something.
My pet theory, is that putting 30c just 'feels' bigger to people. So a long time ago, everyone started putting .30c.
That happens disturbingly often in scientific articles too. Mostly because people don't read many of their references at all.
Compounding the difficulty is that journalists report to non-experts in two ways. Their direct bosses are other journalists. Their indirect bosses are the consumers of media, 99.99% are non-experts. If 99.99% of your customers don't know or care about the details you will naturally put less attention towards the details.
We need to stop pretending that reading/watching/listening to the news is somehow informative or educational. Every piece of journalism is, at some level, fiction trying to masquerade as non-fiction. One way to fix this is to encourage more experts to directly communicate with the public. This has the opposite problem, since most experts are not skilled at communicating their expertise to a wide audience. Ultimately a healthy society requires a balance of the two. I'm firmly in the camp that we currently have too many non-experts trying to communicate expertise.