If anyone would like to ask questions about what we did, I'd be happy to answer them.
There's still lots more interesting stuff to do, but it was enough for a blog post. Suggest away if you think we missed something obvious!
If anyone would like to ask questions about what we did, I'd be happy to answer them.
There's still lots more interesting stuff to do, but it was enough for a blog post. Suggest away if you think we missed something obvious!
Did you try this visualization?: http://bl.ocks.org/mbostock/4062006
We'll have to agree to disagree. I think the visualization you linked to is much harder to read, because the visual weight accorded to each edge is a non-linear (and somewhat arbitrary) function of the 'true' weight. It also doesn't scale well with number of vertices.
Whereas with the chord diagram, your eye is naturally drawn to the big arrows, and you can easily follow them. It's also bidirectional in a more straightforward way.
Also, binning. There is a nice theory for multidimensional binning and aggregation [that I haven't seen anyone describe explicitly so far]. So I wrote primitives. They play nicely with plotting, statistics, etc. That'll also be in Mathematica 10.
Lots of Go for data egress. It's perfect for it.
Also: why Go instead of a JVM lang that can interop directly with Mathematica via JLink?
Finally: Will Mathematica directly support doing the full stack of this kind of work, including the egress?
2. I used Go because I am very productive in Go and like a lot of things about it. Goroutines are neat. Java is fine, it's just very boilerplatey, and I'm not practiced enough at it to get past that. And I don't see why we can't develop a GoLink as well.
3. Probably not the whole stack, at least in the beginning. But we'll get there. We want to make it really easy to spider websites and so on.
I wasted a lot of time trying to do things the "traditional" way by loading into SQL, querying, etc, but it was actually much faster to process things in memory (I have a 16 gig machine). Intensive stuff was parallelized in Go and used ordinary filesystem with directory prefix tries for performance.
Writeup was mostly SW. He's worked on it maybe an afternoon a week for the last month.
I really enjoy visualizations and can iterate extremely fast (e.g. ChordPlot took half an hour). Don't know why M is not the defacto standard for dataviz people. Tweaking takes a long time, and design iterated with me on getting things looking really nice.
All in all, most of my time was spent building tools to easily create multidimensional histograms. The nice thing is that those tools are clearly useful enough we'll integrate them into Mathematica, so the cost is somewhat amortized.
NLP took a few weeks of Etienne's time... once again, amortized. Most of that is wrangling, really, and building tools to understand the deficiencies of your training set. Naive Bayes works surprisingly well, the magic is in the tooling and "human intelligence" you iterate with.
1. I noticed the interactive slider, embedded into the webpage. That's not what vanilla Mma 9 can do. Is there a simple way we can do the same without the CDF plugin (e.g. a package I'm not aware of) or is this future functionality?
2. The graph with the migration data looks nice. Mma 9 can't do an edge layout like this (curved edges) by default. Is this again custom code (custom Graphics or custom EdgeRenderingFunction) or is it future functionality?
3. There's the part with the frequency of various graph motives (the number of edges, triangles, and other weird shapes in graphs). How did you count these? Was it done using Mma? Some of these motifs are easy to count (there are simple expressions in therms of the adjacency matrix), but some others like the (1-2, 2-3, 1-3, 3-4) subgraph are not so easy.
4. How did you make the word clouds? Is the algorithm written in Mma? Here's a nice but slow one (the worse answer by myself): http://mathematica.stackexchange.com/questions/2334/how-to-c...
2. Custom Graphics. We don't (yet) support doing anything interesting with graph weights, which this relies heavily on. I think this is a good candidate for M10 -- name would probably be ChordPlot. I came up with a "GraphForm" construct that allows you to patch various graph properties into various visual parameters (size, color, edge weight, etc). That turns out to be quite useful.
3. Tally[list, IsomorphicGraphQ] . Isn't that cool?
4. Awesome! Nice code. We plan to create ImageCloud and WordCloud functions for M10. WordCloud will be specialized for representing word frequencies and so on. ImageCloud will be the general case: accept a list of images [potentially with transparency], and then find a nice layout given desired sizes. So much cool dataviz will be possible with this! Like country flags...
You can even synthesize C code from Mathematica (there is a symbolic subset of C in it already) and have Mathematica run the appropriate build process for you, so things can get pretty interesting with that alone.
Out-of-core processing of large datasets is already on the roadmap for Mathematica 10. We plan to have a domain-specific language to describe and work with external [or in-memory] datasets in an efficient way, translating as appropriate to the native database query languages. Our 'native' format will be HDF5.
Ultimately, though, I think we'll rely on code generation to compile Mathematica to LLVM or transpile it to Go, so that we can distribute chunks of computation out to a cluster using M as command-and-control.
The idea would be that you can create and test large processing pipelines from inside Mathematica and then distribute them across a cluster in an ad-hoc way, then visualize the progress, track errors, and analyze the results. Notebooks are really good for that kind of lightweight UI.
This isn't a new idea, but in a language as dynamic as Mathematica, I think it could be especially powerful. Of course, it is also tricky because type inference would be a big part of making this idea possible in a dynamically typed, symbolic language like Mathematica. But not impossible, I don't think. And functional languages already have demonstrated advantages in this type of situation -- take stream fusion in Haskell.
What you're saying about distributing a computation on a cluster sounds very interesting. I used Mathematica for a hybrid Mathematica/C++ calculation (LibraryLink) where the complexity was handled by Mathematica and the (simple) heavy lifting by C++. I used the standard parallel tools to run it on a cluster, which means that communication was done through MathLink.
I never went above ~70 CPUs, but people say that problems start to appear above that (too many MathLink connection): http://mathematica.stackexchange.com/questions/20356/mathema...
Another possible problem with my solution (LibraryLink, then Mma parallelization) was that it required a Mma license for as many kernels as I was running, even though most of them were only running the C++ code. But that's easy to fix on WRI's side.
Really nice work.
Some of the plots use this: http://reference.wolfram.com/mathematica/ref/CommunityGraphP...
The underlying community detection uses: http://reference.wolfram.com/mathematica/ref/FindGraphCommun...
If you look under "Method", there are a bunch of different methods to use that I'm told correspond to various landmark papers in the field. If you know about community detection, you'll recognize which methods correspond to which papers, but if you don't, why do you care? At least, that's our philosophy for documentation, but I'm not sure I entirely agree with that philosophy.
https://thestrangeloop.com/news/strange-loop-2012-video-sche...
The naive Bayes and corpus wrangling was done by my colleague Etienne Bernard: http://www.wolframscience.com/summerschool/2012/alumni/berna...
And - is correlations of one's interest vs friends interests the same as correlation of one's interests with itself.
If more fine grained, I wouldn't be surprised to see ties between seemingly exclusive things... e.g. in this http://meta.stackoverflow.com/questions/157976/map-of-all-se... (not interesting, but participation in particular Q&A sites) "christianity", "judaism" and "islam" are in the same category (as opposed to people apathetic to that topic).
It has also made it exceptionally annoying to get feed updates from the 900 bands, actors, movies, songs, etc. that I said I liked and now have to "unlike".
I think you could correlate people within a certain confidence, but because of the nature of the data, you would have to expect a surprisingly large dissimilarity of interests within a clique over a certain account age based on this noise. Not 90%, but higher than the real value.
Do you think analyzing information about social networks will bring out real world insights about people and relationships?
I did some network analysis stuff like this for Twitter long before it was built into Mathematica (rant: someone at Twitter needs to make a 'graph query' API call so that it doesn't take 3 hours to get a single graph of your own network).
I would link to the relevant posts at taliesinb.net, but Posterous is down 7 days ahead of schedule.
I think it can bring out real world insights. You just have to be very cautious and not leap to conclusions because they seem to tell an interesting story. Although it is somewhat disturbing how "gendered" the wall post topic distributions are.
It would be great to play with it a bit.