This coupled with the huge failing of not displaying zero on the y-axis of their graph, and even interpreting the bad graph wrong, makes me not believe them at all. A very low quality article.
This coupled with the huge failing of not displaying zero on the y-axis of their graph, and even interpreting the bad graph wrong, makes me not believe them at all. A very low quality article.
Yeah they interpreted the "toast" graph wrong. They should be more careful to read shitty graphs that cut off at the low point.
i.e. maybe instead of lots of books having direct text like "David said" or "Dora said", over time there was a trend to use a different more varied/descriptive way of describing that, i.e. "David replied" or "Dora retorted"?
To make any sort of claim like "this word's usage changes over time" in an academic sense you'd need to include a discussion of the data sources you used and why those are representative of word usage over time. The fact that they'd even try to use google ngrams in this way shows how little they actually researched the topic.
Google ngrams is a cute data set that can sometimes show rough trends, but it's not some "authoritative source on usage over time" and it doesn't claim to be.
The authors, on the other hand, are claiming to be authoritative and thus the burden of evidence on their claims is far far far higher. I didn't even get into their completely unobjective and vague accusations of "AI" somehow doing something bad. Ngrams don't involve AI, it's simple word counting.
From what I read the authors are only claiming that some Google n-grams fail the common sense test and that the data shouldn't be considered rigorous.
"said" is in the top 300 most frequent English words, according to Wiktionary. For its usage to halve in 80 years then double again in 20 would represent a profound shift in English that would certainly be known to linguists.
Or, as with "toast", one could simply doubt the veracity of the data.
I think this is reasonable. As otherwise we end up accepting things that exist solely, but are flawed. Just because something exists and is easy to use doesn’t mean it’s right.
Just like the answer to “the most tweeted thing is X therefore it is most popular and important” does not require a separate study to find the truth. It’s acceptable just to say “this is a stupid methodology, don’t accept it just because that’s what twitter says.”
This is a reasonable request, but I also think it's fine for the author to state it _as an expert_ that newspapers continued using said at a similar frequency. The story they tell us plausible, and I don't really think the burden of proof is on them.
It's the extraordinary claim that it has that does.
That claim is Google's, and before accusing the author of the blog, maybe how representative their unseen dataset is. Should we take statistics with no knowledge of their input set at face value because "trust Google"?
Instead, we get a strawman put up where they misrepresent what the data set is, make up things that its "claiming," fail to investigate the underlying data sources and look into "why" they see the trend they see, and also fail to provide any alternative data.
It's cheap and snobby grandstanding, ironically complete with faulty interpretations of the little data they DO present.
It should be marked "Fun statistics" with a big red label "Not representative of anything, any graph you see could be and probably is totally bogus" then.
>Instead, we get a strawman put up where they misrepresent what the data set is, make up things that its "claiming," fail to investigate the underlying data sources and look into "why" they see the trend they see, and also fail to provide any alternative data.
A, blame the victim and goalpost moving. Old favorites.
Why the fuck would the author need to "provide alternative data"? Google is showing statistics, that people, including journalists and scholars, take at face value.
Now they're suddenly just "fun statistics", so if they take them seriously, it's on them?
As for why they don't include the evidence in TFA, as others have noted, it's the extraordinary claim that "said" dropped to nearly 1/3 of its peak usage that needs extraordinary evidence backing it up. It's plenty sufficient for them to say "this doesn't make any sense at all on its face, and is most likely due to a major shift in the genre makeup of Google's dataset".