Why scientists need to be better at data visualization
knowablemagazine.org
knowablemagazine.org
From a more recent one ("Simple diagrams of convoluted neural networks", https://medium.com/inbrowserai/simple-diagrams-of-convoluted...):
"[In academic] research, visualization is a mere afterthought (with a few notable exceptions, including the Distill journal, https://distill.pub/). One may argue that developing new algorithms and tuning hyperparameters are Real Science/Engineering™, while the visual presentation is the domain of art and has no value. I couldn’t disagree more! Sure, for computers running a program it does not matter if your code is without indentations and has obscurely named variables. But for people — it does. Academic papers are not a means of discovery — they are a means of communication."
Beyond that, the paper colors reference links #0000ff (neon green). This makes them illegible and unpleasantly distracting. Try a darker (and less intense) color to improve legibility. The bright red links are also a bit distracting, but that’s more a matter of personal preference.
There are some good resources in https://courses.cs.washington.edu/courses/cse512/19sp/ which I just ran across recently. e.g. http://www.personal.psu.edu/cab38/ColorSch/ASApaper.html
For example, in seismology high sonic velocities are always shown in blue/cool colors and low sonic velocities are always shown in red/warm colors. This is non-ideal for colorblind users and confusing for other audiences (red != high value), but it's a near-universal convention. We can use a more perceptually uniform red to blue colormap, but keeping the warm/cool convention is very important.
If your audience is used to seeing a certain type data in a particular way, you'll confuse people if you don't follow that convention. Clear labels and legends are great, but people follow convention first, labels second.
When the goal is clear communication, following established convention is much more important than making a "better" visualization. That's not to say that these guidelines aren't very important, just that it's vital to keep your audience's expectations in mind.
> "Scientists also tend to follow convention when it comes to how they display data, which perpetuates bad practices."
In my experience, scientists follow convention because it enables clear, consistent communication with other scientists in the field. Violating convention is okay, but should be done with utmost care and the knowledge that you'll need to spend more time explaining things.
If I were going to start designing a course in creating scientific figures, I think I'd have a roughly even split between the psychophysics of visual perception (e.g. distinguishing between similar quantities of lengths/angles/colors/etc; designing for color-blind readers; ) and hands-on work in a real programming environment turning data to figures.
That's actually a good reason to learn R and the ggplot2 package. Whenever I write a paper, what I do is that I make a quick shell script that invokes Rscript, with a simple R program that takes a CSV file and outputs a PDF file of the plot, which can be automatically loaded in LaTeX.
Whenever the data changes, it's just a matter of updating the CSV file and running the script that rebuilds the figures and the LaTeX document. As an added bonus, it makes keeping the data with the paper easy, since they're part of the same source control repository.
There are a lot of bullshit techniques used in data viz that are tantamount to lying and people are often sincerely shocked when you call them on it.
"I didn't do that on purpose," as if they didn't learn when they're 4 that it doesn't matter if you meant it if you did it. You still have to apologize and try to make amends.
Things people do from ignorance or malice:
Remove the origin on the graph and the relative height of the lines is skewed. 3d pie charts are 'larger' on the bottom half. Circle charts conflate diameter with area (humans are bad at judging area). On a log scale plot, a fat enough line can make anything look like a trend, because the end of the line is literally orders of magnitude wider than at the origin.
Friends don't let friends use any of these techniques, and the jaded instantly distrust anyone who is using them.
The used bookstore on the closest university campus has an entire shelf full of old editions of Edward Tufte books. I don't know if you'll be so lucky, but it's well worth a shot.
Specifically, if you're really concerned about a delta, removing the origin is good for visualization; it lets you just see the difference between two trends (in particle physics, you can show an energy excess this way). Likewise, for exponential phenomena, log plots are the correct choice, since uncertainty will be magnified for larger values. In both cases, of course, you need to include error bars, but that is always true. But I don't think you can "instantly distrust" someone who is using these techniques when they are valid.
By screwing with the origin you can make the alternative that saw a 6% reduction look much more compelling than the one that achieved a 4% reduction, when in fact it's probably only slightly more compelling.
For example, my lab studies the electrical activity of brain cells. A neuron normally sits at -70 mV relative to the extracellular space. When it fires, it can briefly cross 0 mV (though it probably won’t reach +70 mV) and sometimes you might want to show that complete spike. Other times, subthreshold changes in the cells’ activity are more important to a hypothesis, and then you’ll zoom in on a smaller range (maybe -90 to -55 mV). Neither of these situations would benefit from a plot centered at zero; in fact, it would look decidedly odd.
(Your other comment about log-log plots also seems odd to me, because the width of the line is almost never meaningful, unless there’s something like a confidence/credible interval band, which case that’s the whole point.
Problematic presentation of a similar type is particularly pronounced in, for example, finite-size scaled curves, where the perceived goodness of fit is heavily influenced by the line widths of the curves, the size of the data points, and the resolution of the plot.
These misleading data representations are something that can be found even in recent publications in prestigious journals (I don't think it's always intentional and therefore won't include any references); also, sometimes it's difficult to notice there are issues unless you are very familiar with the methods, or have tried to reproduce the results (and have wasted months or even years of your life by then).
Yes, if the default value is -500, then setting your origin at 0 is going to cause confusion.
I think the problem in part here is that in academia you can make your publishing quota for a period of time on a 2% improvement to some process. In which case you're amplifying the outcomes as a form of self-promotion.
Rarely do software developers, city planners, or who knows what other professions are going to get any accolades or shifts in public opinion off 2% (unless it's taxes). In both cases you're trumping up the numbers. In one it's expected, and I think we can agree that things are perhaps not as they should be in the paper publishing circles without having to open that can of worms to define exactly why or in what ways.
I design my plots for readability and in the process regularly break every rule you listed here, because I trust my audience (other scientists) to understand how to read axes. Yes, it might be confusing to somebody who doesn’t read graphs with axes that often, but if I optimized my papers for their convenience, they would all have to be 100 pages long.
If anything, public documentation is probably a better analogy -- expected audience of developers, but of potentially wildly different experience levels.
And in that case, it's generally fair for say a novice to complain about the lack of examples, clarity,etc.
And these charts are the same -- this is the primary entry point for a novice in the subject, and a helpful tool for experts. A poor chart would be harmful for all expected audiences.
That said, and having looked at the article more carefully (ahem), most of the concrete suggestions they offer are sound, and are things that I think are pretty well-known, at least to the practicing scientists I talk to. "Pie charts are bad", and "Beware of false color images, especially ones that use the 'jet' colormap." are both (parapharases of) statements that I hear from working scientists.
But then there's this: "And yet few scientists take the same amount of care with visuals as they do with generating data or writing about it. The graphs and diagrams that accompany most scientific publications tend to be the last things researchers do, says data visualization scientist Seán O’Donoghue. “Visualization is seen as really just kind of an icing on the cake.”"
This bears zero resemblance to my experience. Most scientists of my acquaintance make the figures first, and then write the text. And certainly I'd say that more care is taken with the figures than the text.
I guess I just feel like I see a certain amount of fetishization of beautiful scientific illustrations (Tufte, etc), out of all proportion to its actual importance to the scientific endeavor. The readers of most scientific articles are probably not going to be flummoxed by a pie chart. And of course, time spent focusing on this kind of stuff is time that might very well be better spent doing experiments.
https://www.amazon.com/How-Charts-Lie-Getting-Information/dp...
Granted, this was sometimes because the only team that had measured a particular substance did it back in the 70s.
It can be quite fiddly, though, so it might not be worth the effort unless you really need it, or you need to do it a lot with many similar figures.
What I would dearly love is some way to animate the systems I simulate. A way to show HOW MUCH flow is going through a pipe, or how much heat transfer is happening in a heat exchanger. Something that's easier to grasp than dull plots.
Sadly, a) I do not even know what to google to find solutions for this, and b) all my primitive searches seem to lead to Blender, which has a large learning curve and way too much time investment requirements.
If you can generate a single static figure, you can manually generate animations by generating each frame separately using your favorite plotting software. You can then stitch them together using many different tools (I use matplotlib) to generate an animation (video, gif, etc.).
What language do you like to use? What's your background -- are you a Matlab person?
Matlab can also do some very nice colormap stuff. Just wish there was a method with less friction!
Edit: OK, there's also gifview.
One of the things we've tried to do -- especially by leveraging the technology and innovations in the tools we build on (especially matplotlib!) -- is to make it easier to make visualizations and plots that are "by-default" communicative and information-rich.
And, of course we have lots of room for improvement, for picking up and encouraging "best practices" at the library level, but it has been really enjoyable and illuminating to have the opportunity to explore that with the community.
Did you consider replacing the rainbow color map images on your front page?
I opened an issue: https://github.com/yt-project/website/issues/70
From what I saw it was C++ code so may not be the easiest thing to write.
Like everyone else I'm super busy but I've been planning on just learning it even though the times I've looked, it's a true blue 90's C++ inheritance obsessed nightmare.
Tbh, for 2D data, I've written scripts that sort of just make a bunch of pngs from matplotlib and string them together using ffmpeg into a movie. For 3D data, I used to use mayavi and ffmpeg to make movies, but that too is pretty clunky. Mayavi feels closest to matplotlib which is why I like it, yt feels like there is way too much boilerplate to just get a plot (for example, I have to specify units for my data before a plot a flat array, I mean wtf).
- Doing research - Supervising students - Teaching - Giving talks in a clear way - Writing papers in a clear way - Writing grants in an enticing way - Performing administrative tasks
This unfortunately is one of the causes why, in certain fields, industry research is much better than academic research. Google is draining brains from compsci departments all over the world, because a scientist at google is mostly a scientist. And there is much, much more teamwork.
Another one is bars in bar/column charts are horribly wide.
This won't work well for a lot of cases. Imagine plotting the atmospheric CO2 concentration time series in the past 200 years. Setting the original point at zero ppm would not make sense because it can never happen. I'd say setting y axis original point at y min is a good trade-off since the plotting package is agnostic about the underlying nature of the plot.
Why does that matter? I disagree. A graph item representing the amount of stuff in some context should double in size (exactly) if the amount of stuff doubles.
But I do realize that practically speaking, it is much easier to set a y-axis to 0 (i.e. a constant value), rather than having to calculate the y-min or near-y-min for every chart.
Do you think that every weather chart is wrong because it doesn't begin from absolute zero?
Usually when charts have misleading y-ranges, the real problem is lack of x-axis context.
For instance, a chart with unemployment over the last three years can be made clearer by showing unemployment over the last 50 years.
Much better to give a wider picture than pointlessly add whitespace.
Based on my own experience, the number of plots that are misleading because they do not start at zero is far greater than that of those that are insufficiently informative because they do start at zero.
That alone is not a reason for "banning" non-zero ymins, but it is definitely a good reason to more critically evaluate their use.