Mining social networks
economist.com
economist.com
Here are few that I know:
For small networks (up to a million or two million nodes such as Wikipedia Link graph from 2009)
Following libraries provide code to handle and manipulate Network datasets:
1: SNAP by Prof. Jure Leskovec [ http://snap.stanford.edu ] written in C++
2: Networkx by Lanl [ http://networkx.lanl.gov/ ] written in Python, esp. good for fast prototyping
There are few Databases for storing networks, e.g. Neo4J http://neo4j.org/ .
Additionally there is a Graph Processing Language called as Gremlin
http://wiki.github.com/tinkerpop/gremlin/ .
For Large networks with millions and billions of nodes, one can use Hadoop / Map-Reduce or Apache Hama [still in nascent stage]. Google has a special system known as Pregel which it uses to perform scalable computations over large networks.
Do you have any resources on how to use these tools (especially the python library), and examples of implementation? I'm interested in learning more and using these tools on real data, but don't necessarily want to spend the time learning all of the theory behind it.
Well you can start with reading this book http://www.cs.cornell.edu/home/kleinber/networks-book/ for overview, regarding application of these techniques are considered you can look for papers at recent WWW, NIPS,ICML conferences. The most popular and well studied areas are Link Prediction and Community Detection. The SNAP library comes with some good example code. You can also have a look at Divisi project at MIT Media Lab if you are interested in reasoning/ analogy over the networks.
For real datasets, there are a lot of encyclopedic datasets such as Wikipedia / DBPedia /Semantic Web/ Music Brainz, as well as social ones such as Twitter follower network dataset. If you are in a university, you can even get full Web Graph from Yahoo [for research use alone].
http://www.schneier.com/blog/archives/2008/10/data_mining_fo...
<quote>But the authors conclude the type of data mining that government bureaucrats would like to do--perhaps inspired by watching too many episodes of the Fox series 24--can't work. "If it were possible to automatically find the digital tracks of terrorists and automatically monitor only the communications of terrorists, public policy choices in this domain would be much simpler. But it is not possible to do so."</quote>
So did something change in the last two years? I mean Palantir is making a crap ton for presumably something similar.
Was this ever a problem, or was it overcome?
The problems Schneier is talking about have very different characteristics than the ones in this article. So the problems Schneier is talking about are still very real. But solving little pieces of the problem is definitely possible, especially with a good amount of human mediation.
As one intelligence analyst once told me " it is looking for a needle in stack of needles"
Whereas palantir situates itself as a tool for intelligence, rather than as an oracle.
For instance, I imagine a high number of updates on a Facebook profile containing the word "drunk" might lead to a higher risk profile for auto insurance.
If you watch some demos on their site you'll see that the workflow is almost entirely manual save for the moments when they retrieve sections of the semantic graph from the datastore.
Also notice that they never talk about the excruciatingly long process of how that data got jammed into the system. Imagine the target goal for an entire day's work for one person is to import 10 news articles into the system and you'll understand why they aren't doing the kinds of things Schneider was talking about (even if their brilliant marketing and sales departments sell it otherwise).
edit: word on the street is that they are still cashflow negative which is very worrisome considering all the fundraising they've been doing + they've almost saturated the government market
Phone companies use this stuff in ways that may be unethical, but it can just as easily be used in ways to empower you to leverage your own network. There are exciting potentials we are just beginning to see.
<sarcasm> How surprising. </sarcasm>
Seriously, it's not like we couldn't have foreseen: http://www.softwarefreedom.org/news/2010/feb/01/freedom-clou...
although sadly: http://en.wikipedia.org/wiki/Hushmail#Controversy
We have a new + faster + easier-to-use webmail coming in the next few months. It's a huge improvement over what we have now.
Search is still in the current folder only though, so that might not be quite what you're looking for.
(btw if anyone wants to help us build a better webmail, we're hiring PHP devs in Vancouver)
Part of their sales-pitch is "Ad-free and Privacy Guaranteed," but eventually it always comes down to a question of trust.