Update on the Twitter Archive at the Library of Congress
blogs.loc.gov
blogs.loc.gov
Great quote in the (too long) nyer article: “Talk, talk, talk: the utter and heartbreaking stupidity of words,” William Faulkner 1927..
I'm sure the NSA, GCHQ, the big platform companies etc are archiving all the tweets forever, it's a shame the public don't have access to curated content taxpayers have paid to keep in those repositories.
The LoB can't possibly archive the entire internet without massive investment
https://blog.twitter.com/engineering/en_us/topics/infrastruc...
Not a lot of specifics, but there's some interesting tidbits:
> Hadoop: We have multiple clusters storing over 500 PB divided in four groups (real time, processing, data warehouse and cold storage). Our biggest cluster is over 10k nodes. We run 150k applications and launch 130M containers per day.
There are a lot of jobs to be had here. But not in tweets. Think cancer research, molecular biology, taxonomy, behavioral and experimental psychology. The data
Did your comment get truncated?
Edit: not to mention that I think you have a far too rosy view of the world we’ll be leaving for people 100 years from now.
The company I worked with spent a non-trivial amount of money storing historical tweets. I'd even go so far to say that was the majority of the IAAS costs - even more than the compute required to process them in real time.
First, how many accounts are we talking about here? Between POTUS, VP, their spokespeople, cabinet members and official agency accounts, I'm going to assume the executive branch has about 50 accounts that ought to be archived. For congress, I think it's reasonable to archive each member's account and their spokesperson: 200 accounts for senate, 870 accounts for congress. In the interest of Fermi estimation (and because I'm not sure if every one of these people has an account) let's call it 1000 accounts.
I'm going to go with a mean of 10 tweets per day (again in the name of Fermi estimation).
With 280 characters + metadata, I'm going to round up to 1KB per tweet.
1000 accounts * 10 tweets/day * 1KB/tweet = ~10MB/day = ~3.65GB/year = ~40GB for the current lifetime of Twitter
If you're drinking from the firehose to archive tweets for a huge userbase (and feed them into models or perform semantic analysis) I could see this getting expensive and costly. If you're just try to archive tweets from a certain group of users and keep them on a disk (with a tarball or zip file released quarterly) it feels a bit more doable.
That said, this little thought exercise has gotten me thinking a lot more about what I expect of the LoC and the National Archives. I'd be happy with a "cold storage" record of the tweets being preserved for posterity, but they may see their mission as making the tweets into a tagged searchable collection. I also think there are arguments to be had about how many people's account really need to be archived (perhaps states could handle archiving their own reps).
In the end I guess I'm cool with the LoC scaling back collection as long as they're transparent about how they're doing it.
You can tweet photos, videos and GIFS. My average gif is 500kb.
If these were excluded I suspect it would make the archive much easier to manage.
The article states that it was archiving _all_ of public twitter. It is moving to the selectivity that you have assumed was status quo.
Perhaps the title should be editorialised to something like:
"Library of Congress announce a change in collections practice for Twitter"
edit:
@dang - how about changing the link to:
https://blogs.loc.gov/loc/2017/12/update-on-the-twitter-arch...
or
https://blogs.loc.gov/loc/files/2017/12/2017dec_twitter_whit...
> “Today, we announce a change in collections practice for Twitter. Effective Jan. 1, 2018, the Library will acquire tweets on a selective basis—similar to our collections of web sites.” The phrasing was elegant, but the sentiment was nonetheless familiar: “Quitting this shit!!!!”
No. That's not at all the sentiment I get from that. The volume has probably grown too big and it makes little sense to archive spend money archiving all of it.
Maybe they're just the voice of perspectives that are more widespread, but every news or pop culture take I see about twitter seems to be born from the unexpectedly-complicated relationship they seem to have with becoming yet another head in an unimaginably large crowd of people. Where everyone can talk to everyone, anyone, or nobody in particular.