Twitpic blocking archive team from backing up pictures
twitter.com
twitter.com
It sucks to see your work crash into the ground like this.
Just so it's clear, there's just Noah (and his parents) that are left running the company.
Dustin Curtis: "Can anyone put me in touch with the founder of Twitpic?(I'll pay to backup all of the photos, and even host them. Ridiculous situation.)"
hi@dustincurtis.com
Just an ugly situation for everyone.
Edit: On second thought, I agree that it's probably related to the bandwidth cost.
I wonder if the Archive Team has considered sending twitpic a hard drive and asking them to just rsync the data over?
Edit: This assumes that they keep backups of said data outside of EC2. Walking a hard drive over to your local Amazon DC.. won't quite work.
Dustin Curtis: "Can anyone put me in touch with the founder of Twitpic?(I'll pay to backup all of the photos, and even host them. Ridiculous situation.)"
Let's not pretend that Noah is down to his last packet of ramen noodles, here. Twitpic is a long way from being a "failed company" by any reasonable definition of the term.
"Archive Team has 900 IPs aimed at Twitpic. We're grabbing 100 pictures a second."
He subsequently corrected this figure to 42,000 photos per hour, which works out to about 11/sec.
Later he adds:
"In some cases, twitpic (before removing all the image access) was banning entire ISPs to stop archive team backing it up."
Maybe they just don't have $15K to pay for this. (and that's not counting all of the users who are also downloading their data)
People have been trying to talk to him since the day the announcement was made.
The cost for an export would be 1/10th that, and would be covered by others, if only Noah Everett were cooperating.
We're talking about 800 MILLION pictures here. Noah is apparently too busy to save them, but not so busy that he can't find the time to actively prevent others from saving them. That's pretty hard to support.
When you abuse the public trust, you get flak for it. That's a feature, not a bug.
No wonder they failed at business if this is how they act during hard times.
I have serious doubts that here have been any actual Twitpic employees for a while. If there were, that seems unwise of both them and Everett. Twitter announced the doom of third-party photo sites in late 2010. Twitpic basically went silent a couple years ago. Meanwhile, Everett says he's been working on Pingly for over a year[0]. Note the timing of that article -- not even a month before the announcement of Twitpic shutting down.
No matter how much he wants to try to make this about a trademark dispute, it's clearly not some sudden collapse. By all appearances, he's moved on to other things and has finally realized it's time to shut down an old project that's been on a long and probably inevitable downward spiral. This is understandable, failing to preserve data is not.
[0] http://www.postandcourier.com/article/20140810/PC0313/140819...
I note, by the way, that the arguments have degenerated from reasoned to "but rights!!!". You're going to have to do better than that.
Edit: While at the same time twitpic had no announcement on the frontpage and users were still uploading: https://twitter.com/textfiles/status/522835537543974913
http://twitpic.com/account/settings
I wish I could say this will be a good lesson in not trusting startups with personal data, but it won't be.
"t" is a Base64-encoded blob of data containing a few miscellanous bytes, as well as the URL the proxy should download (in this case, twitpic.com/show/large/...). "s" is the signature, used so we can't just abuse the proxy to download more.
However, as this is the internet, until news breaks otherwise I will continue to believe that Hanlon's razor applies here...
Never attribute to malice that which is adequately explained by stupidity.
They'll probably be able to start new things without most people noticing. And they probably know this.
Second, even if people didn't understand that, the HN crowd is a tiny slice of the good hackers out there. Most of the best people I know do not read it, or even actively disdain it. "Cesspit" and the like are a terms I have heard in person or on newsgroups used to describe the comments section here.
Third, I'll bet most HN people don't even know the names of Twitpic's founders. I don't. I wouldn't be able to identify whether another company, down the line, was founded by them.
Lastly, I doubt most HN people are so principled that a year or two from now they would turn down an otherwise compelling job offer because the founders of the company were once associated with banning Archive Team's IPs. It's mildly annoying at worst.
This could impact a launch significantly, but of course there's no guarantee the masses would notice.
Anyone boycotting twitch.tv? Anyone????
We very well could be having the same situation in a few years when Twitter is in the process of turning off the servers.
If you want to keep your data, make backups. Don't trust cloud providers to be there. If you're the Internet Archive, make backups/dumps/archives before people start turning the servers off.
The only thing is, it's empty. I don't know how you get a 14MB zip file that's empty, but Windows is telling me it's empty. That's not cool.
14:23 <REDACTED> you have to enable the "show hidden folders" in windows, under their view options
14:24 <REDACTED> each version of windows is slightly different but it is usually press alt, select tools, folder options, view, and the radio next to show hidden files and folders and then ok
14:24 <REDACTED> then you close and reopen the zip file
14:30 <REDACTED> oh, and sometimes the first download is corrupt.
14:30 <REDACTED> I've had to help too many people at work with this problem plus a person or two on this channel
Of course hosting the application is a whole different thing, but hey, at least you got all your data available.
It would be like the world's biggest art gallery turning around and saying 'Oh, sorry, they won't let us call our gallery McDonalds, so we're shutting down and burning all your pictures. After all, easy come, easy go.'
At some point we have to learn this lesson for good, and stop trusting random entities to be good stewards of things we find meaningful.
The primary pain point when spidering it was that each GeoCities user had a tiny hourly bandwidth cap; to do the job right the spiders had to keep track of the error responses that they got so that they could go back later in the hope of getting the real content rather than the error message.
It strikes me that a similar fate would most likely have occurred if Everett was hit by a bus and his bank accounts eventually ran dry.
We should identify historically significant single points-of-failure and make archives before the failures occur, insofar as is possible.
If all of twitpic's data is exported or copied elsewhere, he can no longer effectively profit off of it. We don't know that his intention is to just nuke the data and be done, he may be trying to salvage what he can monetarily (and the "we're getting acquired, just kidding" thing is probably good evidence that something like this is going on behind the scenes). I don't think there's anything wrong with this either. It's his service, and if his motive was profit, it makes perfect sense to shop the data around instead of just accepting a total loss. There's no reason twitpic should have to be run for altruistic purposes.
Maybe dcurtis et al would have better luck getting a response if instead of saying "We'll cover the bandwidth costs", they said "We'll buy the data from you at a price commensurate with the years of effort and upkeep you've invested into an apparently massively important historical archive". I don't think it's fair to pretend like Noah Everett is obligated to sell his archive for the cost of bandwidth.
In any case, at this point he's made it clear that he wants to wash his hands of the project in one way or another. It's his personal project and he is and should be free to do that. This is the worst possible time to try to do a bulk archive.
I understand that we can't see into the future and can't always predict when something like this is going to go down, but I'd think we can probably improve our heuristics so that something as seemingly significant as twitpic can't just vanish into the ether next time.
Eight hundred million photos and all their comments. That's historical record. That's our historical record. No one jerk gets to destroy all of that. The "hurr, well, don't use it" is completely bonkers--I don't use it, but my history is on there, too. It's a record I didn't put there but matters to me as a citizen and as a member of society. Destroying a major part of the historical record for everybody is indefensibly shitty, full-stop, you don't get to say "well it's mine" when it's not.
It's ours. Twitpic was allowed to profit by holding onto it awhile. Now it's time to do the right thing.
If I were him, I'd probably be contacting Sotheby's or another auctioneer that specializes in high-value goods and seeing about selling off the dataset to the highest bidder. Then these fellows can put their money where their mouth is and show us just how important the preservation of that data is.
b. If it's worth so much, people should be willing to pay a fair price for it. The implication that he should let it all go for a payment to Amazon is, frankly, insulting, and I'm sure you'd be singing a different tune if it was your own startup that you were trying to liquidate.
c. Talk is cheap. It's easy to try to shame someone into letting you take their intellectual property. It's not plutocracy, it's just business.
d. Getting a lot of users doesn't require your company to convert into a charity. I am 100% sure that Noah Everett has a price, and I'm sure if people actually valued the Twitpic corpus he'd be able to get it.
e. All uses of Twitpic were voluntary and done with the implicit acknowledgement that the images and comments may go away one day. If someone was not OK with this, they should've saved their data, and many probably did. Blocking the Internet Archive from devaluing the dataset is not bad, wrong, or ethically icky in any way. The larger ethical crime lies at the feet of the people trying to convince Everett that he doesn't deserve fair compensation for curating this historically significant dataset.
f. If no one is willing to pay the money to preserve the dataset, it's obviously not as culturally significant as you believe it is. There's plenty of money out there.
- The whole shutdown / acquisition / shutdown saga - Tepid response wrt getting compensated to export data
Someone's still holding out hope for an injection of capital. Because once the export happens and the data is gone, the value is entirely lost.
I wouldn't bank on getting any traction on paid exports until the last possible second.
It might be a better path for people to reach out to help the guy WITHOUT offering to offload the value proposition of what he built.
Also, any sort of agreements between people who uploaded content and Twitpic can't just be sold off to the first person to show up with a hard drive or offer to pay bandwidth costs. That would be a set of legal costs, etc. etc.
I can't imagine they'll want to spin up new machines, so it makes sense to throttle/limit/block.
Not sure I'd immediately jump to malice here.
I'm not saying that it isn't malicious. I just don't think it's wise to jump to that conclusion straight away.
Good luck.
If there are 100 pieces, and 500 people, you want each piece to be held by 5 people for 5x redundancy. And, if nodes leave the swarm, the system automatically rebalances which nodes are maintaining a particular piece of the dataset.
The only solution being worked on out there that does this is IPFS
https://github.com/jbenet/ipfs https://github.com/jbenet/node-ipfs https://github.com/jbenet/go-ipfs
My bet is that this is just a giant media whoring play. If you want to save the company data then you have to be transparent and ASK for help. If it is going to cost $15K someone might very well step up. But if you are tightlipped then people will assume the worst.
They cannot do it any other way - the Internet Archive does not have infinite resources to host things, and they need to prioritize. 120TB is not an easy thing to host.
The Archive Team is the "oh shit this site is shutting down and it's not all in the Internet Archive already, we gotta save it before the deadline when it's gone forever" squad.