MTV news website goes dark, archives pulled offline
variety.com
variety.com
Journalism is unique in that it's almost always public in some form. It should be a reasonable expectation that it stays in its original medium, or is accessible in archives if that medium vanishes. Major newspapers often offer reprints or back issues. NYT offers the "Times Machine" [0] with basically everything they've ever run digitally. This should be the standard, not the exception.
In the more common case where you are not, though, the law already provides exceptions to copyright for archiving and redistributing materials that are no longer available. This generally only applies to content that isn't audio/video/graphical, but a wider exception is provided specifically for "news" (journalism).
If you want to negotiate a contract making such archival exceptions explicit within the contract itself, that's probably not a bad idea. Even if not though, the law provides for archival protections already.
Further reading: https://www.law.cornell.edu/uscode/text/17/108
I think you hit the nail here.
Most contemporary journalism is just not worth archiving. There is a small percentage that definitely is worth archiving, but the bulk of it is produced for the moment.
Even NASA (accidentally?) recycled the master tapes containing the moon landing AFAIK, and some of them are ultimately recovered.
Shortsightedness, I may say.
I don't agree. There are many interesting tidbits which go unnoticed as bog-standard / disposable news until the dots connect in the future, and one wants to construct the history backwards from that tipping point.
When these "contemporary articles" go missing, many important details of the history is lost.
One of the people who recovered a lot of old Usenet archives once remarked that the reality was that no one cares about the details of some long ago SunOS bug whereas a lot of the cultural discussions about government policies etc. provide a window into the time.
Almost all “journalism” today is opinion notifications without substance, source, or first hand accounts. I suppose that’s its own kind of historical artifact, but not the kind I’ve seen be meaningful in my study of history.
History isn't just about the big things. The history of the mundane is also important in understanding a culture.
The issue is we don’t know what may be important after many years, decades or centuries have passed. Given how easily text compresses, it would be a shame to not have at least text of real news sources archived perpetually. I know there is also an endless stream of listicles which are just generated to get clicks and ad views, but newspapers are worth archiving.
I do see your point. I'm reminded of the photograph who snapped basically a digital throw away picture of Monica Lewinsky weeks/months/years before anyone really knew who she was. later on, he was happy to have that picture, since it was one (or the only one) of her at some event hugging Clinton.
Text is easy enough to store, but making it useful to search and access seems like another problem to solve.
Maybe the comment thread on a trivial clickbait article contains a post by a future dictator, and it will be a crucial piece of her biography fifty years from now.
But to a historian it is all a series of moments. To understand how/why one happened, it is useful to know the ones that came before, and things created for the moment might better describe it than those with a wider purview.
Of course you don't need it all, but deciding what to keep is a complex (and bias ridden, accidental or otherwise) problem in its own right.
Considering how much historians love ancient rubbish bins, one mans junk is another mans treasure. Even the junk produced during and before the Trump/Brexit 2016 election/referendum would be massively valuable to a future historian in explaining how those things happened.
If we allocate 200KB per article for pretty strongly compressed images, then it's still 100 million on a single hard drive. Why not archive the whole lot and let God or future historians sort it out?
You only need to start getting picky when video is involved, but on the other hand when the alternative is total obliteration you can crush video down to 1MB per minute and have a tolerable VHS-like experience. And even 8MB Shrek gets the point across.
That's still pushing a terabyte a day[0] on YouTube alone. What about TikTok as well?
[0] https://www.statista.com/statistics/259477/hours-of-video-up...
There's some logistics in handling backup tapes, but that's a well understood process.
https://routenote.com/blog/wp-content/uploads/2020/08/pex-yo...
And going by this, restricting to a reasonable view count will cut the space you need by a factor of 10 to 25.
If we round that to a petabyte per year, it's not in the range of a typical personal archive, but it's a reasonable thing to picture several big libraries doing.
(Though a single person could handle that much if they really wanted to, spending $10 a day on data tapes.)
We could train a model every year and preserve it, then future historians can quiz this model that thinks it’s 2024 and ask it whatever they need to know. It’s fascinating because it will probably “know” the kinds of everyday normal people things that are very hard to glean from only reading old news stories. Things like how the average person feels about their world, or how they feel about current events and why.
It's published at virtually no cost compared to what newspapers and magazines used to cost. And with LLMs it's going to get even cheaper as you no longer need people to write the stories. We are spiraling into an era where we will be drowned in content that is all worth next to nothing.
It is impossible to know presently, how valuable something will be to someone in the future. I don't know if any historians have this as their motto, but I like to think they do.
Laws were passed to update this to the digital age, allowing the deposit libraries to archive UK websites. Unfortunately they don't right now archive video. It's also a bit limited given that web publishing is not national, so much UK content is published on US run websites.
https://www.bac-lac.gc.ca/eng/services/isbn-canada/Pages/cre...
Perhaps younger readers have never experienced looking at microfiche of old newspapers in a library. Analog machines, not computers, but, if memory serves me correctly, the archives were far more complete than the so-called "tech" company-mediated, advertising-based nonsense what we have today on the www. And the analog fidelity was great.
Google ultimately failed to "organise the world's information" to the extent that it previously was organised by public libraries. Instead they turned finding public information into a game designed to support a massive programmatic advertising racket. The www has more volume perhaps than pre-internet libraries but 1. an enormous amount of it is garbage and 2. it is not organised for systematic browsing (leading to discovery) nor serious, methodical research. It is optimised for "clicks" and "views", selling ad services, not learning. There is no way to "browse the stacks" to see how this information is organised. It's all "secret". Imagine a public library that had to hide its operations as a "trade secret". Imagine that it tried to intermedite which stacks a patron will browse, for commercial purposes.
The point is that public libraries had extensive collections of newpapers and periodicals. The public could do serious research. But the www is dominated by advertisers and most of the information published is influenced by commercial motives. And it's all intermediated by a gatekeeper that sells ad services: Google.
If I want to use the internet (cf. www) for library research I still somtimes search library catalogs via z39.50 if the library's website is cludgy.
Man, I miss the original Google mission.
Who would you trust more? Big Tech and Silicon Valley? Or a non profit with a quirky benevolent dictator whose mission is universal access to all knowledge?
The entire point is to make the Internet Archive have a legal team that can tell crybully journalists and their uncles to eat a fat dick
Yes. It's very easy to get into the mode of "I can always download it"--until you can't. Or at least it's very hard to find. I've been pretty good over time. I probably have copies of 75%+ (probably more for stuff I actually care about). But I've definitely written for sites that don't exist any longer, at least a couple of which were behind paywalls.
Not just digitally. Times Machine has everything going all the way back to the first issue of the New York Times on September 18, 1851: https://timesmachine.nytimes.com/timesmachine/1851/09/18/iss...
That's like reading this HN thread in the year 2197.
But the mining is optimized for storing as big of a fraction of the chain as possible, so I worry that the future of it devolves into a few big nodes and everyone else dropping out.
And maybe take legal action against those who've already used it, for training and knowledge bases, without licensing it.
That'd be different than a company simply not wanting to incur the small costs of keeping it online. (Still sounds crappy, but it's not "for nothing".)
I think one of the worst crimes of modernity is to lock away knowledge and art for the sake of „protecting and sustaining it“ - copyright is broken and needs reform, the current system does not serve it’s intended purposes and we need to come up with something better.
If you preserve it, provide access and have the ownership rights, you keep it. If you don’t, you lose it.
For example, there was a large classic cartoon archive which had a whole team of retouchers painting out all the smoking.
I have a few of these: Nirvana, Joy Division, Pink Floyd, Uncut Guide to Shoegaze. For $15 or so each, it's not a bad deal, about the cost of a CD. Full-size colour pictures look way better than the low res images found online too.
They are clearly fine associating the brand with the most bottom bog reality shows as long as the views come in and the budget is non-existant. Wouldn't be the first brand Viacom/Paramount gutted.
So not all, only "older" articles are not available. Wayback Machine only works after someone submits the site for archival; the old site might have used something that made it hard to archive content; and I expect MTV pre-dates the Internet Archive. Any of those factors might result in WM missing stuff. It's not magic.
I don't know for sure, but, as a hypothetical example, in this case, MTV could have requested that IA take down the archives from that time period.
(That is, if MTV can prove they owned the domain during the archive period [probably pretty easy]), and if they still hold the copyright to that work [probably], then they could essentially issue a DMCA takedown request [by email], with which IA would most likely comply [there are some HN and IA forum posts about this I think].)
> just because the parent/ source website was shut down
The "just because" part is not a well-founded assumption AFAICT. For all we know, there could have been a request like above.
> why wayback machine and the other archive services aren’t filling in this gap/working in this case?
I see no evidence that IA isn't working in this case, if working is defined as including having a process for handling valid takedown requests that must be acted on according to the rules they follow.
I hope that helps. Again I don't know if this is what happened, it's all hypothetical.
The analogy that comes to mine in regards to MTV taking their full site off-line, would be the equivalent of a newspaper shutting down and destroying or burning all of the extra printed copies of their paper/articles they may have in their own archives instead of donating them to a local library.
We were promised 21st Century jetpacks and all we get are same old dated mindsets, dated biz models, etc.
p.s. While we're on the subject, anyone want to recommend a Firefox extension that does full page capture (read: not a screen shot)? And then a simple in-browser DB for saving / cataloguing with tags or similar?
Or even credits at a solo site level. It's not about time but articles reads (i.e., actual usage).
Also, ending my subscription shouldn't end my access, just access to new content (i.e., content I did not pay for). And maybe the older content gets ads as a supplemental.
My beef with the 20 yr old article is... it's been openly available... and it's 20 yrs old. Now I need to sign in to see it? That's not a paywall. At this point, that's just stupid.
https://www.nngroup.com/articles/the-case-for-micropayments/
Again, even at the solo site level it's not being done. Certainly, it could be.
I think you'll find though that building a wallet is an extraordinary complex problem. The legal compliance issues alone make it tough to build a service that can operate in the USA. As soon as you build a service for transferring money, criminals will immediately try to use it for money laundering and other illegal payments. What's your plan to handle that? You can't just ignore it or most likely the government will shut you down. And those AML/KYC laws are unlikely to be loosened just to suit this use case.
Why? Why would I start a business where there is no market? We've gone over this. It's not a technical and/or solution issue. It's that the gate keepers - the ones who would but the product - don't want something that's good for consumers but bad for them.
I can't explain it aby other way.
So if i want to read one article from website XYZ, i may need to pay 00001 credits and if I want to read another article from website yyy, i may need to pay 00002 credits?
How would yyy or xyz content creators/journalists actually make a living from this, though?
Content is a meaningless filler word, and I don't care if "content" dies.
I use Single File.
> Save an entire web page—including images and styling—as a single HTML file.
In any case, if that's the capture, what's a viable long term pro-privacy storage option? Private GitLab or GitHub? Maybe use Issue somehow as more of DB?
Probably best the jetpacks aren't made by slaves, so I guess they need paying for.
In the current operating environment, it is important to optimize for optionality.
(No affiliation with the Internet Archive, just a concerned citizen)
The LoC seems to have some web archiving going on, but it’s very selective. I don’t know how any good archivist can think they will know what will be important in the future.
I did a search for MTV and only got 3 results, and only 1 is actually from MTV. For the cultural significance MTV had, I find this hard to believe.
https://www.loc.gov/web-archives/?q=MTV
I can’t figure out how to actually view the archived page, which is also an issue. What good is an archive if it can’t be accessed.
Either way, it seems like if the government was going to put more money behind internet archiving, they’d fund their own service to make it better.
Don't take it for granted.
I mean, when it dies it will be a greater tragedy than the burning of the library of Alexandria (at least many of the books there were copied elsewhere), it will be a tremendous blow to our collective knowlege and memory.
But it's inevitable that it will go away.
In the long run we are all dead, yes. But whether it will be around in the next decades is up to us who are currently alive.
I think its a fair way to have some check and balance between a centralized library doing all the archiving, and a million libraries each with their own archival system.
Edit: and this case drives it home. Copyright makes it illegal to preserve this music history, and the owners don't give a shit.
Also a mandatory CC/NC licensing for some type of content could be a good idea
It's not illegal to archive of course. Video recorders were a thing even back in the day of MTV, completely legal consumer devices.
I think saying that duration isn't a factor is oversimplifying the issue. We'd be relying on multiple independent archivists to all hold individually-obtained copies of the data for nearly 100 years, and none of them would be allowed to transfer the data to anyone else. That's a very different scenario to one where the data were freely mirrorable after 24 years.
Edit: it's not the archiving that's illegal (probably)--it's the sharing of the archive that is. And sure they can share it after the copyright term runs out, but that's so far in the future the odds of an unmirrored archive surviving are slim.
Same goes for series that have been altered to avoid music royalties [2]. Pirated version will remain as released.
[1] https://en.wikipedia.org/wiki/Harmy%27s_Despecialized_Editio... [2] https://www.fastcompany.com/91109690/why-streaming-platforms...
Of course, this exists: <https://en.wikipedia.org/w/index.php?title=Harmy%27s_Despeci...>
But, to another point, copyright is a convenient bogeyman in these discussions--and it plays a role. But someone would also need to pay for archiving all this material and indexing it in a manner that would actually be useful.
More specifically, Greedo didn't shoot at all. Not until it was messed with many years later.
HN discussion to an article about her from a few days ago: https://news.ycombinator.com/item?id=40702546
I'd be concerned if the archives of, say, Kerrang! looked at risk of being lost. Or even back episodes of Pimp My Ride or Celebrity Deathmatch. But I didn't even know MTV had a news.