BBC To Delete 172 Websites Due to Budget Cuts, Geek Saves Them for $3.99
readwriteweb.com
readwriteweb.com
The cost to the BBC to keep those pages on its servers is more than $3.99. Just because one guy downloaded all of the pages, compressed them, and then seeded them on bittorrent for a cost of $3.99 to him doesn't mean jack.
Notice that he's not hosting the content in any easily readable form. No, he decided to put that burden on everyone by putting it up on bittorrent. Why isn't he hosting the content? Because hosting a heavily trafficked site ain't cheap.
People were going to lose access to that information, and now they aren't. That's really all that matters. With torrents, the burden is on "everyone" only if they really want to keep it up.
edit: Reasoning for downvotes would be appreciated. The BBC has a long history of recklessly losing valuable data. See kgtm's link to previous yc post on this topic.
edit #2: Oh, okay, I see what's going on. When I said "means jack", I didn't mean to imply that the BBC should spend inordinate amounts to keep the data up. What I meant was: Yes, it might have cost them to do it, but look, like what this guy did, there are other ways to keep the data archived. Hell, I'm sure some people would even settle for a CD archive or something.
Actually it doesn't mean jack. that's my money they'd be using, and from a quick scan of those websites, i'm happy for them to disappear.
For some reason, many people seem to have decided that all data is important and there's an almost fetishistic devotion to saving everything that can be saved. Not all information needs to be preserved for all time.
I disagree, and strongly. Future historians will probably disagree with you too. Even the most inane TV shows can prove to be extremely valuable when trying to decipher the culture of various countries 300 or 3000 years in the past. Imagine reading a trashy romance novel from 1000 BC. You'd learn a lot about the people who lived in that time.
You might counter this by saying that with all the information out there now, we only need to keep the "good stuff". But who decides what's good, right now? How do you know that what you decide is good will always be seen as good? Your personal opinions on the content disappearing are totally irrelevant when perhaps millions of people have seen it or have been affected by it.
All information that is created should be saved, especially since we're able to. Imagine the ancient Romans deciding to burn a bunch of books because it cost too much to hire guards for the library. Ouch.
Most of these sites are promo fluff for BBC TV shows. Some are cancelled, some are for news show with little more "content" than "This show next comes on --- on BBC -".
Most of these sites would be dismissed as "marketing" if the BBC were a for-profit company. Marketing for defunct products. The rest is material that's being legitimately archived elsewhere.
There are teams of people right now deciding what is 'good' (i.e. what is kept) and what is shredded. They're called archivists, and, using their judgement in conjunction with government-created records retention schedules, they're culling both government and private records. This culled information is removed for a variety of reasons, space being one of the most prominent. By such means is the pool of available primary sources that will form the basis for future histories created by underpaid and underfunded but well-meaning bureaucrats.
Sounds like the same situation obtains not only in the dead-tree archives but also in the digital ones. Sigh. I wish people would prioritize our cultural legacy first instead of last when it comes to funding.
If these sites were a special instance of our culture, then fine. But they're not. We don't need every single episode of Eastenders archived for all time to understand culture. A good sample is way better.
These are different times to the classical era. We have a glut of information, most of which, once archived, will never be looked at again, and will only cost money until finally abandoned.
I'm sympathetic to preserving the shows themselves.
Their promotional pages on the BBC site when they're no longer in first-run? Feh.
Speak for yourself!
As far as losing valuable data, well okay, they maybe blanked a few copies of Doctor Who here and there, but it was not considered valuable at the time, and as always cost was a factor then as it is now, it's easy to pass judgement with hindsight.
I upvoted you because I agree about the content, so I can only speculate about downvotes. If I had to guess, I'd say it's probably because you either missed or ignored the point you replied to, which is that the reporting was, indeed, terrible.
This is quite a fascinating topic in itself. http://en.wikipedia.org/wiki/Category:Lost_BBC_episodes http://en.wikipedia.org/wiki/Doctor_Who_missing_episodes
Rehosting it isn't legal either, but then neither is torrenting it.
I've seen previews of suggestions for legislation but no actual legislation for this sort of scenario, can you point me to some.
Making money from something isn't a requirement for copyright infringement, nor even is it the test for commercial use of copyright works.
I've never heard of a "fair use" case on historical (preservation) grounds can you point to one? As for educational use in Europe one generally doesn't have the relatively liberal educational use allowances that one does in the USA, and from hazy memory I think that the USA legislation does tie down quite well what is educational use. Certainly unlicensed redistribution except on the most de minimis basis wouldn't get a pass as fair use for education.
IANA(IP)L and am a little out of touch wrt the latest legislative efforts.
---
OT rant:
Just recently it's interesting but all my posts in which I've tried to appraise people of copyright legislation or at least the potential threat of such have been modded down substantially. I have pretty strong views contrary to the current established laws myself, FWIW, but I feel that even if you're an anarchic dissident it's important to know the laws one is fighting against. Clearly others here disagree that reminders about demands of IP legislation are worthwhile for this type of forum.
Since you seem like you may be interested: looked at http://www.groklaw.net before? They've had quite a bit of activity on IP laws, last time I checked (as has the rest of the internet, but still).
Note that large portions of geocities are being hosted on the web by various parties, most of whom are apparently using pizza money to do it. Geocities was still in the top-100 websites or so when it closed, and they're probably getting decent amounts of traffic from people trying to find old sites.
Even if there is heavy traffic, you're hosting a archival copy, not a production site, so you can just degrade performance as needed. See often slooow web.archive.org :)
No, he decided to put that burden on everyone by putting it up on bittorrent.
Give a man a fish and he can feed his family for a day. Let the man catch his own fish and he can feed his family and community for a lifetime.
Also, just because the info isn't public does not mean the BBC doesn't have it stored in their own digital archives.
(BTW, did anyone notice that the BBC calls these sites "TLDs", and that they took a beating for that by people who assumed they were incorrectly referring to "Top Level Domains"? But, TLD can mean "Top Level Directory" too, which is entirely accurate. Most amusing...)
Anyway, I think it's good that we're developing a culture where some people care, and make sure that things get archived, even if it's done sometimes unnecessarily, and almost always only gets a degraded copy (ie, a web rip without streaming videos, and without structured data). I only hope that these vigilante actions don't lead companies to not pre-announce massive data erasures.
Having downloaded it, it's around 2GB compressed. It contains images, but most pages are nonfunctional, due to links being specified from the site root /. To view it properly you need to place the folders at the root of your web server or of your hard drive.
The full list is in the torrent file.
[1] sambeau, http://news.ycombinator.com/item?id=2188870
Also, keeping these sites (without actively updating them, monitoring comments etc) would essentially cost the BBC nothing, especially the low-traffic ones. BBC Online is a very lean organisation.
Why is this a story, even in web circles? Why all the hyped-up outrage?
Are people associated with the BBC shilling this?
I downloaded one of the sites in the torrent titled Zombies. It's about a British girl who organizes a community effort to make a Zombie Movie. Sure the site was ripped but not the video content. Nothing on the page is worth seeing other than the video. Once the BBC turns off access to that stream the archive is virtually useless. I'm going to checkout more ripped sites; my gut says they're probably video heavy too.
I'd say that most of the valuable content is still stored on the BBC servers, in the DB where it's being served from right now. I'm sure they're going to repurpose the good stuff if they haven't already. I have an acquaintance who works in the web department at the CBC. They've implemented a pretty neat CMS to repurpose their older media to work with their latest site. I can see the BBC doing the same.
Further more, many of those assets are accessed via iPlayer which has restrictions on access outside of the UK.
Even if that is the case, I'd sure like to see a breakdown of expense when a cut like this is made.
From what I've seen reported, the only site of any significance which is being removed is 'WW2 People's War', which asked people to put their and their family's personal stories of the second world war into a central archive (http://www.bbc.co.uk/ww2peopleswar/). I believe that this was already being archived by the British Library, so won't be lost.
The overall decision to cut 50% of the 'TLDs' is definitely political, to placate an administration that's beholden to the Murdoch media empire, but the choice of which ones to kill seems mostly pragmatic.
Intuitively, it seems that the answer should be 'not much' -- perhaps on the order of a few thousand dollars/year. A couple of servers, bandwidth, and a sysadmin checking in now and again to apply security patches &c.
But here (also with e.g. yahoo closing geocities), it's argued that the cost of keeping them up is much, much higher. Where does the expense come from?
The author: they really should have known better than "s/he", to boldy write for RWW.
Edit: Added a sourcewith examples - http://www.crossmyt.com/hc/linghebr/austheir.html and the Wikipedia page seems to have some discussion about it too - http://en.wikipedia.org/wiki/Singular_they
"I will say of that person: they laughs a lot."
(Compare: "I will say of that person: she laughs a lot.")
"They is the person you should talk to."
(Compare: "He is the person you should talk to.")
Or should it instead be: "I will say of that person: they laugh a lot."
"They are the person you should talk to."
? And why?However, in the general sense I would argue that the conjugation of verbs follows a pattern but that pattern doesn't have to be singular vs. plural. Given "he has" or "she has" one might expect a singular "they" to follow the pattern "they has". However, "I" is singular yet it uses "I have", so it's not unexpected for the singular "they" to use "they have". In other words, the conjugation of "they" is the same regardless of whether it's singular or plural.
Correct: "Grammatically, this paragraph is incorrect. It is grammatically incorrect. I'm sick of these grammatically-incorrect paragraphs!"