I wonder if the answer is a network of topic-focused archives; like moving from a "Library of Alexandria" model to a modern nationwide system of libraries.
It raises the issue of governance of the curator, but the IA is already more transparent than Goole & co.
The nation of, say, Japan, has limited interest in funding an american noprofit today; but they would likely have a great deal of interest in funding an equivalent focused on Japanese content, for example.
So now you get into the issue of haves and have nots. Who is allowed to be considered an authorized archivist from a robots.txt perspective? Or what happens if an archivist becomes blacklisted for not respectfully crawling? How do national sanctions affect the Internet Archive of Russia? I imagine there would be a certification process and it would probably cost some money.
It's an interesting topic and I'm simply looking at the weak spots. I'm not against the overall concept though.
Distributed governance on the internet is a massive issue, and it's effectively unsolved for everything from pairing to DNS. In practice, good faith goes a long way, particularly in areas that are largely academic in scope - like archiving.
If most of it is crap I would call not archiving it a feature.
There is a weird convoluted analogue to CERN particle detectors. They smash particles together and then image the resulting storm of particle contrails via detector that is basically a sandwhiched ccd detector (like you have in camera, but different) the size of a cathedral. Resulting in far too much data for any system to analyze or even store in the first place. Hence they need/needed to runtime filter the massive amount of particle trail signals and only pick out the critical ones.
If there is too much data you simply need to drop the parts you are fairly confident you don’t need.
There is no reason there should be only one internet archive, there might very well be parallel operations filtering a bit different things.
I guess it’s a bit odd Unesco does not already have a parallel effort.
The internet becomes so full of hallucinating AI output that it becomes impossible to train the model.
(heh - checking that link Charlie Stross just posted (Jan 31) a blog post: "An AI app walks into a writers room")
Anyways the link I was actually after ... give https://www.antipope.org/charlie/blog-static/fiction/acceler... a read and consider the "what happens in the later parts of the book."
I share the same hope but have doubts that as a society we'll have the collective critical thinking skills to disconnect from the AI overlords. We've already had the US and US inspired Brazilian coup attempts fueld by social media placements and it's only going to get more fine tuned and effective.
What can I do as an individual? One path is to simplify and declutter my digital life. How else to cope?
> The dark forest theory of the web points to the increasingly life-like but life-less state of being online.Dark Forest Theory of the Internet by Yancey Strickler Most open and publicly available spaces on the web are overrun with bots, advertisers, trolls, data scrapers, clickbait, keyword-stuffing “content creators,” and algorithmically manipulated junk.
> It's like a dark forest that seems eerily devoid of human life – all the living creatures are hidden beneath the ground or up in trees. If they reveal themselves, they risk being attacked by automated predators.
> Humans who want to engage in informal, unoptimised, personal interactions have to hide in closed spaces like invite-only Slack channels, Discord groups, email newsletters, small-scale blogs, and digital gardens. Or make themselves illegible and algorithmically incoherent in public venues.
Unless I'm completely misunderstanding how ML works, which very well may be true.
This will not work, its too soon.
https://i.imgur.com/B2cHXRA.jpg
To me this implies that things such as "sarcasm" is a pattern simple enough for an AI to match - and that should go both ways, whether it's being generated or recognized.
If you're arguing that it won't be able to detect the more subtle sarcasm, then yeah, sure. But, well, Poe's Law predates GPT.
No, you got it right. Describing it as a game of telephone is a great analogy. This is exacerbated by the confidently incorrect problem. LLM output looks sophisticated and correct and may at time actually be correct. However some unpredictable percent of the time it will be incorrect and confidently so.
I suspect we will see the rise of both groups of machines, curated A.I.s and A.I.s just trained on anything, which should be entertaining.
https://i.imgur.com/u8Np332.png
The curation, such as it is, appears to be limited to humans downweighing the undesirable answers. Which is why there's always a way to work around it, even though it requires more and more elaborate prompts.
Isn't that extrapolating the current trend a bit too much? Clearly, the text corpora[0] amassed before mass LLM content distribution are already big enough to train such models to decent general language fluency. So why would AI creators contaminate those datasets with potentially spurious content?
Sure, you want to keep your model up-to-date about the state of the world (the GPT corpus ends in mid-2021 afaik), but you can be much more careful about which texts you include. Those newer training data serve a different purpose than the original corpus, you don't need to bootstrap general language proficiency anymore. OpenAI already released a product for classifying AI-generated text, why would they not use something like that to filter future training data, for example?
[0] edited, thanks!
Man I'm so tired of this very obvious observation. I wouldn't think a company smart enough to create an AI would also be dumb enough to fall into a pitfall that even the most casual observer can identify.
And I'm not looking forward to a swathe of AI generated songs pumped into the charts and streaming services at potentially lots of songs per second.
In the Kessler Syndrome analogy, that's an ablative aerospace impact armour company deliberately launching and blowing up satellites to sell their goods to spacecraft builders.
——
We will pay you 100 bucks to withhold filing your partners death certificate for one week and providing their certificates to us.
Meta collapses if all its properties are filled with AI generated spam content.
Same for Google.
Meta will likely fall later since visual content at scale is still 12-18 months away. But for Google, the clock is ticking.
I don't want to read 10 pages of text for a simple cooking recipe.
Its completely empty calories in terms of knowledge.
Begun the Bot Wars have.
Although I hope we may see big come back of Web 1.0 forums such where users have to gain street cred and even invite referals with realy genuine contribution into community, no way to fake with AI today.
Stable Diffusion provides already good enough images to cover a lot of visual content online.
If you look at image shares sites for the purpose of entertainment like Imgur, you will also notice that a large portion of the viral content are screenshots from Twitter or traditional media.
Is content opinions? Most people don't have opinions on every topic on the planet. Is Gobekli Tepe the place of Noah's Ark? What's going on with Hunter Biden's laptop? How will Meta's VR strategy work out? Will it rain tomorrow in Sydney, Australia? Depending on your area of interest you might or might not have an opinion about it which you may or may not publish online.
Divide up the 'net into trusted and untrusted sources. Make the trust ratings public. Use search tools and corpuses such as the Google Books dataset to source "knowledge" back to pre-Internet roots, when necessary. In short: bring academic reputation back and bring it back hard.
It will make for a more elitist web, but given that even without ChatGPT we've had a problem with wildfire misinformation spread in social media networks it might be a change that's a long time coming.
What it means is the bar for becoming a new StackOverflow contributor (or Reddit admin, or Wikipedian) might become much, much higher. "Oh, you want to contribute your first post? Show me the bicycles in this image, find the letters in this image, and provide the names of two existing Stack Overflow users with over 1000 karma who can vouch for you, and also you see a tortoise on its back, baking in the sun. You're not helping it. Why aren't you helping it?..."