Stuff that every programmer should know: Data Visualization
c0de517e.blogspot.com
c0de517e.blogspot.com
But if you need interactive 3-D plots (e.g., a wire frame that you can rotate), look elsewhere.
That's why unfortunately I still recommend something like processing, then you can write your own small IPC to send data to your visualization code from any host
Unfortunately, to my knowledge there aren't any comprehensive textbooks that cover visualization from the ground up. We didn't use a single textbook in my 2-year degree; all lectures were heavily based on research papers. Central topics if you want to read up on this is perception (which color scales should you use? how many parameters can you plausibly put in one plot?), different visualization techniques for different data (scatterplots, histograms, treemaps, horizon graphs, volume rendering, graph drawing with edge bundling, +++), interactivity and applications of basic techniques (Visual Analytics, Interactive Visual Analysis).
A multitude of scientific fields use different visualization tools, so it can be tricky to find the relevant material for whatever it is you're working with. But in general, I think the data mining/big data/analytics fields could do very well with a bigger focus on visual techniques. If you get the right visualizations for your data, the truth often just jumps out of the screen. GPUs can let you work with multi-gigabyte datasets at interactive framerates, although I haven't seen a lot of practical applications of this yet. Can also be used for non-spatial data, if you're clever with CUDA or just use the shader data structures creatively. Would be interesting to hear if anyone in the industry uses this yet.
You can use server side javascript which should handle 20GB dataset without problems.
I use d3 on the client with a 80GB (currently) dataset, by putting the dataset in elasticsearch. It's a pretty fantastic combination. You can do multi-value aggregation from unstructured data, or geo-spacial searches, or lightning-quick full text search.
The server has 8GB of ram and 2 cores, and with about 1.2 million new documents every hour, barely breaks a sweat.
At the moment I roll up daily stats and store them in a separate database for longitudinal analysis, but eventually I'd like to ship data that is more than a couple of weeks old to hadoop.
I've added a few (well many) links at the bottom of the article but if you have any suggestions on resources/software/etc please let me know
It felt like my typical high school class. They'd teach us how to calculate the circumference of a circle, but they never told us what we'd use it for. "Programming" is not specific enough.
One of the things I've picked up more recently is to take a presentation or whatnot and mentally block all but the top 10-20%. If I can't get the bird's eye summary of what each slide needs to say there, then my manager won't understand the story. They're {busy|lazy} and don't care to read the entire slide - that's all chaff for the underlings to digest so they know the finer details.
So for a programmer, one of the biggest uses of good visualisation skills is the selling of ideas. A good plot goes a long way in convincing someone of something.
I've included a few pictures of stuff I recently used, maybe it can help to give more context.
- http://1.bp.blogspot.com/-YA10ftXFFQ0/U6-iAdAUtGI/AAAAAAAAAn...
This one was done for debugging, as I do often. The code that I was debugging took some geometry and generated some texturemaps from it, it's really hard to debug why the maps are wrong when they are, what happened to the geometrical calculations. So I just added a std::vector<float> debugStuff and pushed values from various locations, then wrote it as a CSV. In processing I load this bunch of floats I know the meaning of and plot them in 3d as a point cloud. Each point in the cloud is clickable to show more of what happened at that position.
- http://1.bp.blogspot.com/-EFgdBffKEU0/U6-h9QweujI/AAAAAAAAAn...
This was for performance. A particular system computes tens of thousands of generated code snippets according to some rules, that will then used in runtime in various ways. Testing in runtime is hard because it's hard to cover all of them and it would take a long time. So I rigged the system to extract statistics about the generated code and save them to a CSV. Then in Mathematica I categorize and plot these datapoints and I can compare what changed between two different runs of the system. When I see in the data that a given change seems promising enough, I do the expensive test in runtime
Any production environment data visualization is going to run into a plethora of sticky problems. How do ensure your queries aren't going to overload and crash your visualization client. How do you handle time series and gaps in data? How do you evict data from a vis?
Some nice irony too: There's a typo in Jonathan Schwabish's bio with visualizaing.org instead of visualizing.org
And yes ... I can easily imagine flogging exploratory gloves and goggles to impress the Board and let them surf through data looking for insights