Newswire: A large-scale structured database of a century of historical news
arxiv.org
arxiv.org
The entire concept of an "NDA" has been bastardized in this manner, if you think about it. Conceptually you may think of an NDA as protecting sensitive data from disclosure, a sort of intellectual property right. However, it's been co-opted by folks to do nothing more than cover up inconvenient truths because they realize most people cannot afford to either (a) give up money they are promised in the future or (b) bankrupt themselves in their own defense.
So it's basically a game where whomever has the most money can ensure their narrative wins out in the end, because competing narratives can simply be "bought out".
I understand a lot of other outlets have been folding and making deals with OpenAI, because they're too weak to sue and desperate for the revenue.
Which if true is really sad. It's like taking out a payday loan: solve your short-term problem by giving yourself a bigger one in the future.
:-)
Makes me wonder if this is a bigger problem.
The FAQ at the end is a bit confusing, because it says:
Was the “raw” data saved in addition to the
preprocessed/cleaned/labeled data (e.g., to support
unanticipated future uses)? If so, please provide a link or
other access point to the “raw” data.
All data is in the dataset.(^- oops, example transcription error)
The "end of described process" dataset descrption is at https://huggingface.co/datasets/dell-research-harvard/newswi...
and for each record there's "newspaper_metadata", "year", "date", and "article" fields that link back to _a_ LoC newspaper scan.
I stress _a_ singular as much is made of their process to identify articles with multiple reprints and multiple scans across multiple newspapers as these repititions of content (to a degree mitigated by local sub editors) with varying layouts are used to robust the conversion of the scans.
I haven't investigated whether every duplicate article has a seperate record to a distinct scan source ..
So, what items of news cause the most "change" (or surprise) to that group of people. All our understandings and notions of truth are based off of large chains of reasoning that are based off of premises and values. When a new event happens that changes a premise our understandings are based on, which events cause the largest changes in those dags?
We're regularly subjected to "news" that doesn't change anything very much, while subtle events of deep impact are regularly missed. Maybe it would be a way to surface those things. I wonder if someone smarter than me could analyze a data set such as this and come up with a revised set of headlines and articles over the years that do the best job of communicating the most important changes.
If it was all powered by AI I think you could get some really interesting results. I bet it would invoke extremely strong emotions in Humans, they tend to not like their "reality" being messed with (well...in ways other than they have become accustomed to - these ways they seem extremely fond of, and defend them passionately).
Another angle no one's run with in any serious way would be a sort of meta-journalism, again with the techniques described above. I think a well done implementation could steal/borrow 50% of Trump's base, and 30% of the Democrats. Actually, with different variations of parameters I'd think you could get a wide range of outcomes, it is often hard to know in advance what will strike a chord with people, some of the weirdest things work like a charm.
TLDR: it makes sense :)
https://play.clickhouse.com/play?user=play#U0VMRUNUIHllYXIsI...
> For belter safekecping Russta’s $2¢4,000,000 collection of crown jewels, probably (he finesl array of gems ever assem- bled at one tle
https://play.clickhouse.com/play?user=play#U0VMRUNUICogRlJPT...
May also be easy to correct a lot of it:
“For better safekeeping, Russia’s $24,000,000 collection of crown jewels, probably the finest array of gems ever assembled at one time,”
I want original text, including misspellings, and original regional / historical spellings, including slang (which may look like another word, but is not, and isn't in a dictionary).
You cannot fix OCR text wirhout lioking at the original.
vs willingly knowing you are introducing corrections that are ridiculously wrong.
Advocating and being a champion for inaccuracy, really isn't a positive. You should find a new thing to quote about yourself.
And thanks by the way for the readiness to jump to conclusions and fire a salve of allegations, viz. "willingly", "knowingly", "introducing", "ridiculous"
AI is a ridiculous answer, with its hallucinations and absurd error rates. If you didn't intend to support that level of absurd error rate, you shouldn't be replying in defence.
It sounds like you did not want to give that impression, if so, I suggest you look at the chain of replies, and the context.
AI hype is literally a danger to us all.
> Prime Minister of the United Kingdom, from 1940 to 1945 during the Second World War, and 1951 to 1955
> Died 24 January 1965
It's so nice and refreshing to see something like this, instead of the common "we tweaked this and that thing and got better results on this and that benchmark."
Thank you and congratulations to the authors!