I've stopped using box plots (2021)
nightingaledvs.com
nightingaledvs.com
I feel like this is where the confusion stems from for the author and everyone else here. Box plots don't make anything bell shaped (they don't change the distribution), they assume that your data follows a bell/gaussian shape. This is correct in cases where the central limit theorem can be applied (which is almost everywhere) - but when that is not the case, the assumption is wrong and you shouldn't use a box plot anyways, because the values it shows have no real use. There are very real use cases for box plots, but people need to understand the basics of statistics before they can use them.
Edited to preempt nitpick.
Because real data is never perfectly Guassian, or perfectly anything.
But the idea of a box plot is that it's for data which is in theory Gaussian or a similar unimodal kind of bell-shaped curve.
Then you can look at the box plot and see if it actually is -- are the two boxes roughly equal-sized? Are the lines a bit longer than the boxes but not insanely so?
No, as the sidethread comment notes, there is only one way you can compute quartiles. You seem to be arguing that the correct thing to do is to impute them, and that calculating them is such a deviant practice that it would need to be specially remarked on.
Box plots are made for visualizing generalized normal distributions and nothing else.
And now people in this thread argue you can calculate them from something else. Not sure if you are replying to the right post.Your theory would imply, among other things, that the median line going through the box part of a box plot always divides it in half, which obviously is not the case.
Whatever you do, you should explain first what you do that your whiskers stay meaningful and are not just whatever randomness your outliers produced.
I agree that if there is an indication that if most professionals don't really know what boxplot is supposed communicate, maybe it should not be used.
Any set of numbers I give you, you can compute quartiles for it. There is no algorithm for doing that that breaks down if the numbers don't follow a normal distribution.
When you calculate the box plot using normal distribution parameters, the outliers are outside the outer bracket.
If you split the dataset into 4 equal parts, the bracket will be larger because the outliers are still inside it.
The methodologies are not equal.
This thread is the first time i heard people do the "split dataset into 4 quarters" and using that for box plots.
In any event, none of these methods assume normality, or rely on CDFs of a normal curve.
If they did, every box plot would be symmetric.
The fact some people think that boxplots are constructed in such a way is a pretty good reason to take the author's article seriously as for how boxplots are confusing.
It serves to distance it from the moment-based statistics like mean and variance at least.
The SVG you've provided clearly shows that the box plot splits the data in 4. The interquartile range (IQR) is clearly marked and it even has a comparison for what the standard deviation (variance) measure would be.
Secondly, if the data truly came from a normal distribution, there are no outliers. Outliers are data points which cannot be explained by the model and need to be removed. Unless you have a good reason to exclude the data points they should be included. This is why I like the IQR and the median, they are not swayed by a few wide valued data points. The 1.5*IQR rejection filter I think is lazy and unjustified. Happy to discuss this point further as it is a bug bear of mine.
What you want to explain to me (IMHO to the wrong person) is the correct approach of calculating a mean and standard deviation and drawing the box from that. Lets stay with that (and thats what i said earlier in the thread)
After i wrote the post you replied to, i realized that the pure "splitting" method for box plots is nonsensical since the outer brackets interval is determined by the two most extreme values. They are too random to be meaningful. It does not make sense to draw a box plot from that.
If you want to represent the standard deviation with your box plot, you can calculate it using standard formulas, many maths libraries have them built in. I don't know how to plot it using any graphing package though. ggplot, plotly and matlab all use the quantiles (the ones I have experience with). Perhaps where ever you learned to read them as mean and standard devation has a reference you could use?
> They are too random to be meaningful. It does not make sense to draw a box plot from that.
This can be a problem. In practice, the distributions I see don't go too crazy and are bounded (production rates can't be negative and can't be infinite). I prefer to use the 10th and 90th percentiles which are well defined and better behaved for most distributions. I do make sure it's very clearly marked on each plot though as it's not standard. Using the 1.5 x IQR cutoff is no better though as when you have enough samples you find that the whiskers just travel out to the cutoff.
The data has been reduced to three numbers, throwing away most of the information that you would need to assess whether the distribution is gaussian or not. If it's not, how will you ever know?
People waltz in with assumptions and then complain when they don’t work because they don’t really understand the tools they are using. The author is one of them. It’s a bad article and the author should not be using or demonstrating things they clearly don't understand.
If it's a bad plot, perhaps some introspection is required...
Not sure how to square that with this statement on Wikipedia's page on box plots:
Box plots are non-parametric: they display variation in samples of a statistical population without making any assumptions of the underlying statistical distribution[3]
My point, again, was: just because a boxplot is not useful to some people, doesn't mean that it is not a useful plot (particularly when augmented with a rugplot or a strip plot). Plots are not just used to convey information to others: they are also a useful tool in exploratory data analysis.
Notice that you can also apply the same critique to almost any plot: some people don't know how to interpret a violin plot (or kernel density estimate plot) correctly... does that make them useless?
The main advantage of a boxplot is that it is parameter-free (unlike histograms, violin plots and kernel density plots) and quickly conveys very specific information (median, range, quantiles, confidence interval for the median) that other types of plot usually don't.
Quantiles and medians. (Plus min and max.) Non-parametric.
I've never understood this to be the purpose of a boxplot, only a means of visualizing a distribution's quartiles.
You've gotten a flood of comments from upset people, so I'll keep it short by saying that a boxplot doesn't actually do what you claim for Gaussians, as the 0 and 100 percentile "whiskers" would be at plus/minus infinity. As for a bounded bell-shaped distribution, there are several non-unique ways to define such a distribution.
The point is not to plot an ideal Gaussian, the point is to plot the data.
In real life the whiskers are the actual minimum and maximum values observed.
Look at this: https://upload.wikimedia.org/wikipedia/commons/1/1a/Boxplot_...
0.7% of all values are outside the whiskers.
The very Wikipedia article your image comes from explains this:
I think this is a misunderstanding, and I think it is shared by the author of the article. Boxpolots show ranges. That's it.
The author's example has a bimodal distribution (TWO peaks) and chooses a type of chart that has ONE peak (a box plot).
A little baffling tbh.
And a histogram for the author's example is perfectly acceptable to show that single data series.
But imagine if you have 10 different normal data series and you want to compare their medians and distributions between each other... well are you going to put 10 histograms side by side and expect the reader to compare them? No -- that's where the box and whisker plot shines.
The author's first histogram clearly shows most of the distribution lies in [20,100), then the [10,20) bin is empty but the [0,10) bin is quite full. Hence, that's not a single-mode distribution. It has two modes, one around [50,60) and the other in [0,10).
To my mind, if you have a genuine EDA attitude you plot it all.
Well no, because you can compare the datasets by eye and say questionable qualitative things about them, but you can't make definitively true quantitative statements about them.
Show me two plots of data points and I can show you two people who will in good faith argue over which one has the higher mean or higher median or higher variance. Because you often can't tell.
The entire point of something like a box plot is that it does part of the quantitative analysis for you. You can see where the median is. You can see the width of the quartiles.
CDF plots are great for plotting a single distributions, but contain way too much information if you want to plot 6 distributions next to each other for easy comparison.
Violin plots are interesting but also quite complicated, since you have to arbitrarily choose a kernel shape and this artificial smoothing can make it look like you have much more data than you really do.
I really don't like the author's "alternative designs" because I think they're even more open to misinterpretation than box plots. It's hard to judge though, because the central problem is that the author is trying to represent a bimodal distribution, and shouldn't be using box plots or the 2 "alternative designs" for that.
Huh? A box plot doesn't have any peaks. A box plot is a histogram subject to the constraint that every bar in the histogram is equally tall. There can never be more than zero peaks.
The author argues otherwise, can you give an example of a use case where box plots would be preferable to the alternatives the author suggests?
> The author argues otherwise
No, in the article he says he wouldn't recommend them _in most_ situations. It's a part that a lot of people here seemed to have missed whether arguing for or against box plots.
>Despite making more visual sense than box plots, I still wouldn’t recommend these design concepts or box plots in most situations because…
(Emphasis mine)
„So, no, I can’t think of any situations when a box plot would be the truly best choice, other than those in which the audience demands box plots because that’s what they’re used to seeing. If you can think of any such situations, though, please let me know on LinkedIn or Twitter.“
„Other reviewers suggested that the conclusion should be that box plots are a useful chart type, but only for statistically savvy audiences. Again, I’m going a step further, suggesting that even those audiences would be better served by other chart types in virtually all situations.“
> Design concepts such as the ones below make more ‘visual sense’ than box plots:
present the exact same info in much less visually confusing ways, through the use of brightness (weight) and area. Just better box plots.
And of course you can always draw some lines for the quartiles on any kind of plot with a linear scale for the value.
That IS a bell curve. While it's true that the Guassian distribution is often called a bell curve or even "the" bell curve, a non-Guassian single mode distribution is still absolutely bell shaped in a general sense.
So, although you started your comment with "nah", you're actually in agreement with the content you replied to.
You could include a lot of little bells far from the single mode, but that's reading a little too much into the literal meaning of "single mode" - a "bimodal" distribution isn't one where the two most common values are both modes. It's one where there are two distinct local maxima.
The tails to the left and right must asymptotically approach zero (or you don't have a smooth distribution, because you have discontinuities somewhere), and if there's just one local maximum, your curve will look like a bell.
[1] https://en.wikipedia.org/wiki/Gamma_distribution or https://stats.libretexts.org/Bookshelves/Probability_Theory/...
The bottom whisker contains 25% of the data, yet is just a thin line, which can furthermore be arbitrarily short.
It really is a dumb visual presentation.
The only way to use it is to recover the five parameters from it, and then stop looking at it.
For that purpose, a QR code would be just as good, if not better. You'd need a device with a camera to get the parameters (but "everyone" has that now), and when you're looking at it with your bare eyes, it doesn't tell you any visual lie.
...which is its intended use case since Tukey invented it as a way of visualising the "5 number summary". I think part of his criteria were that it should be easy to make by hand which is clearly no longer a consideration so there are plenty of reasons to just do something else most of the time these days.
No they don't. They show quartiles mostly, and don't assume symmetry or any parameters of a gaussian.
Quartiles are not relevant, i.e. can be highly misleading, for a bimodal distribution or beyond...
But even in the first image of the article the fact that two quartiles are close together means that there some density peak around there.
I agree with the author that box plots are not good plots, but quartiles/deciles/medians are useful even for multimodal distributions
> they assume that your data follows a bell/gaussian shape. This is correct in cases where the central limit theorem can be applied (which is almost everywhere)
You sir just failed basic statistics
As far as I can tell you’re making the introductory student error of thinking the central limit theorem means any sufficiently large sample makes a distribution look normal.
Most natural distributions are not Gaussian upon sampling even if they’re bell shaped. They often have fatter tails, model some complex process, etc. The box plot is sometimes deceptive as is demonstrated in the original link. I don’t think that’s easy to argue against as they provide a totally reasonable and common sample distribution and show it failing to be descriptive of the most important features.
I fail to see how the CLT even remotely addresses the concerns or obviates them in any useful sense. The CLT and box plots aren’t very often applied together for these reasons.
A lot of people here are commenting that no, technically box plots don't assume any distribution. And I mean, technically you can ride from NYC to SF in a lawnmower.
But I completely agree that box plots shouldn't ever be used for anything but unimodal distributions similar enough to a bell/gaussian distribution.
All of the criticism of the article seems to be that they're misleading when the distribution is not bell/gaussian, e.g. bimodal.
To which my reply is, of course. Box plots shouldn't be used then. But if your distribution is bell/gaussian, they seem fine and I see no particular issue with them.
Boxplots are a single tool for data analysis. They do not apply in every situation, nor do any other tools. The same goes for pie charts, which are constantly being accused of always distorting data. Pie charts, like box plots, have their place.
If a plot can mislead, it will.
Or take the first example from wikipedia page on box plot [0]: "Box plot of data from the Michelson experiment", which is just 20 points per run. Would I want to see this in the paper? No please. There is no evidence that the experimental data is gaussian (or even single-modal). Or further down that page, "A series of hourly temperatures" - why would one box-plot it either?
And even if you claim your data is gaussian by construction, maybe because you surveyed lots of people - I still want to see the evidence, as it's pretty simple to make experimental mistakes that turns data non-gaussian (say you only surveyed two neighborhoods with very different properties)
In other words, the domain where box plots are sufficient is very small. Most publications should never use them.
This is debatable, but noting that box plots are satisfactory for unimodal gaussian-ish distributions is not a very persuasive response.
Additionally, I find the article informative but believe it could be improved with this clarification. As someone who has worked with data analytics but is not a mathematician or actuary, I know people who probably review these types of graphs. Now, I understand that it is essential to check the underlying data distribution to avoid being misled by the information, even if the source and axes seem trustworthy
I think the word you're trying to use is "bimodal" and yes, that is one example where the author's reasoning fails. But it's not the only one.
>I couldn’t find any papers or information on this.
You said you have no formal higher education in mathematics - how would you even go about finding (let alone understanding) papers? Regardless, just to be clear, this is not something you would learn from papers but from introductory textbooks and university courses. Everyone who has to deal with statistics in science needs to go through a whole lot of extra education exactly because there are many pitfalls like this.
>it is essential to check the underlying data distribution to avoid being misled by the information
That is another half-truth that everyone on the outside seems to agree on, but it is useless in practice. What do you do if the underlying data is not accessible. And what if you don't have the means to process it for every paper you read (which is what usually happens)? Then you have to rely on the actual tricks of the trade, which will come naturally if you worked with tons of statistics before. There are lots of telltale signs that let you spot bad analyses by only looking at a plot or summary chart. Granted, you won't catch all of them, but it often takes real malice and deep statistical competence on the author's side to cover up these things.
A box plot isn't trying to show the same thing a histogram is, it's like saying we should stop using Venn diagrams because they confuse people when trying to show the exact amount of overlap, so pie charts are better...
It's silly.
Not true. Box plots represent low and and high quantiles independently and support the representation of outliers.
This alone is enough to make it clear for everyone that they don't require a distribution to be symmetric, let alone bell shape.
Violin plots and bee swarm plots are better. Jittered strip plots can be okay if you're careful to avoid saturation (or more points added in the saturated region will disappear as they can't make it any darker).
You can even make 'em show histograms: https://miro.medium.com/v2/1*J3Q4JKXa9WwJHtNaXRu-kQ.jpeg
Here is a great rant (borderline lecture) from Angela Collier on why they aren’t [0]
I think it's useful to be able to compare the approximate shapes of histograms during exploratory data analysis. Is the thesis of this criticism that this isn't actually a useful thing to do, or that violin plots don't achieve this, or is it "just" an aesthetic argument?
EDIT: Maybe she'd be fine with using them in an exploratory manner. She seems to mainly be complaining about using them in publications, meant for other people to consume. Also: I did not watch the entire video (:
Compared to these alternatives, violin plots are comically bad.
They look like vulvas. We're all adults, it's not a problem typically, but given that it's an aesthetic choice (noticing how half of the chart conveys the same info without this property), why? And it does come up, like if someone does make a joke about it, a room full of typically only well-meaning men will now look to her if she's comfortable with the joke and, what was okay before, now turns into a feeling of being singled out and outside the rest of the group
1) To show the distribution, in which case just the histogram arranged horizontally in the traditional fashion is far better than a violin plot with 2 copies of the histogram vertically and some extra quartile stuff tacked on, especially since lots of standard libraries to do violin plots do kde with very extreme smoothing so the distribution they show can be very misleading as to the real empirical distribution.
2) To highlight the summary statistics (quartiles and median) in which case just the boxplot is better because generally these are hard to read on a violin plot
In case #1 this is usually because the distribution differs significantly from a Gaussian in some interesting way that would make a boxplot irrelevant or misleading. (eg it is bimodal or multimodal).
In case #2 this is usually because the distribution is Gaussian (or otherwise standard) and you want to compare it with other standard distributions. You don't need all the information in the histogram and to include it all would obscure the important point(s) you're trying to make about the median and quartiles. What is considered standard is going to depend a lot on the domain, audience and subject matter. In her case, she's an astrophysicist, so if you're looking at say red shift data from some observation, other astrophysicists will know the distribution you would expect to get from that sort of observation for example.
That video is basically a summary of all the conversation attached to this article in some ways.
In any case - I don't personally use them not because of that but because of the reasons I gave[1] which she also mentions in the video - you usually want to present either the distribution (in which case a horizontal histogram without extreme kde smoothing or quartile info is usually better) or you want to highlight just the summary stats in which case the boxplot on its own (or just a table) is generally better. When I find I want to call out a given summary stat (median/mode/some quantile cutoff) on a histogram it's usually better in my view to just show the cutoff on the histogram and shade the tail (eg you frequently see hypothesis tests as a histogram with the critical region shaded and the CV1 number or whatever called out specifically).
[1] and one other which is they are even more confusing in many respects for non-experts than a boxplot so if I was to put one in a presentation or whatever I would find myself spending an undue amount of time explaining the plot rather than making whatever point I wanted to make with the plot which is never a good sign. It would be different for someone who tends to write for/present to fellow experts I imagine.
And I just don't relate to this at all:
> you usually want to present either the distribution (in which case a horizontal histogram without extreme kde smoothing or quartile info is usually better)
Where I almost always see this is in time series plots where there is a distribution at each point. Horizontal histograms are not as intuitive for visualizing this, because plotting time on the x-axis is so universal. And while it is true that box plots work well for this when the distribution at each point is close to normal, it is not true that all data looks like this, and it's easy to not notice this if you default to using a box plot.
I do agree with this:
> or you want to highlight just the summary stats in which case the boxplot on its own (or just a table) is generally better
Yes, but you can also just leave off the summary stats from the "violin plot" (just like, as you point out, histograms usually don't and shouldn't include summary stats) in order to visualize only the shape of each distribution.
I also really don't care about the flourish of vertically centering / "reflecting" the distribution, a series of vertical histograms totally expresses the same information that I'm saying is useful here! People seem to find that ugly, which I figure is why they started doing the reflection thing to make it prettier, but I really don't have a strong view either way on which of these presentations is or isn't ugly or leads to awkward jokes. I just think "a series of distribution shapes laid out vertically" is a commonly useful visualization.
And I really don't know about your last point; I don't spend much time working with non-experts who don't understand histograms really well.
Is there a different name for the version of this that doesn't include the summary statistics on the same graph? I think seeing the distributions at different x-axis values (in my work, nearly always in a time series), but including the summary statistics is not as important and I agree that it's noisy.
Similarly, the distribution represented by a box plot itself is often the distribution of "just one sample". When viewed as such, a distro has its own uncertainty[1] and that uncertainty is not represented in a violin plot, for example. As with every "right tool for the job" debate, people will vary based on experience with the tools, including how to simplify/explain them to others.
What I don't see is anyone saying "box plots are useful because they're the best kind of chart for [specific use case]". I can't off-hand think of any situation where I'd rather see a box plot than a strip plot or violin plot. When and why would you want to summarise the data so coarsely and visualize it so un-intuitively?
Box plots are useful because they're the best kind of chart for when I have multiple populations and I want to quickly glance whether it's reasonable to assume that the populations have the same median, or not (you do that comparing not just the medians of the populations but also the shaded areas)
Sure, the jitter plot provides more data, but if you only make use of the quartiles anyway, that extra data is but an unnecessary distraction.
One solution is to smooth into a kde and then use transparency to indicate overlap, but that’s introducing more complexity than you want for a quick n dirty first pass
How is the reader to know you've used the right plot? How are they to know that you haven't hidden a bimodel dataset behind a box plot because it makes your conclusions easier?
> If the distribution is more complicated and you need some detail, use a histogram or a ridge plot. Violin plots are never the best option; they're curvy so a little more pretty but don't do a good job of conveying information.
They are just multiple, non-overlapping histograms plotted next to each other. They allow you to compare distributions without them getting in the way of each other.
I can understand if it's the fitted PDF that you think hides the original data. That is unnecessary IMHO.
Important caveats: the generating processes for all these quantities are the same in a physical sense, so they are comparable. All the distributions are roughly lognormal-ish, so they are single-peaked distributions, as folks are discussing here. The point of the visualization in theses cases is not to understand the properties of the distribution per se, it's to show the important percentiles because they have business implications.
The point is that they are parameters of relevance to observer.
I work in medicine sometimes work with box-plots for this reason. The questions "what is the 25th percentile outcome" is perfectly legitimate
Sometimes less is more; Box plots are specifically good for showing and comparing quartiles.
If you want to compare several groups and care about gross differences, they are an excellent tool. They are an excellent to when you believe the data is normal and think the histogram is misleading. they are also great if you think the data isnt normal but care about quartiles.
Any time you would be happy with a table of the 5 datapoints (min, max, median, 25th, and 75th percentiles), box plots a great tool for graphic comparison.
But after having read these comments, it really drives home his point that you can get a room full of lots of very smart people who all know what they're talking about, and they'll all disagree about the understanding and interpretation of box plots.
It's a little surprising, but the evidence in these threads pretty much cinches the argument for me.
They are mostly useful for comparing batches not analysing an individual batch.
The author doesn’t know what they are talking about and is telling people as if they do. If he read any of Tukey’s material he might know. But no name dropping is enough clearly…
Yes, the author is aware of that. They even stated so:
> Despite making more visual sense than box plots, I still wouldn’t recommend these design concepts or box plots in most situations because…
Seems a few people missed the "in most situations" part. He's saying he stopped using them for whatever reasons because it isn't working for his audience. So as the title suggests, maybe we should all take a look at our use of box plots and see if there are better alternatives.
Also remember who he's talking about when it comes to reading box plots. He's not talking about people who understand box plots. He's talking about others that don't know or understand box plots, which seems to be thousands of people he's had to explain it to, according to him.
Not only that, the cases presented are likely better dealt with via inference tests. But the author's knowledge doesn't extend that far. And even going as far left, the posed question isn't even defined in the article. So how was a suitable methodology chosen? Well it wasn't - lets just throw this pretty picture up and whine about it.
The author is way out of their depth and should retract the article and take a formal, accredited statistics course.
I'm not one to appeal to authority but "author should take a course" is akin to ad hominem when a quick look at their profile (https://www.practicalreporting.com/about-nick-desbarats - https://www.linkedin.com/in/nickdesbarats/) tells you that he's been doing dataviz and statistics for a long time.
I'm not one to appeal to authority either which is why I am making objective arguments about what is presented.
And yes he should go on a stats course. I dread to think the chaos he’s spread to people who don’t know better.
Ad hominem arguments, however, should also be ignored. Saying that I'm unqualified doesn't prove anything and adds nothing to the discussion.
If you have specific criticisms of my reasoning, I'm more than happy to listen. If all you have are personal insults, however, well, enjoy the rest of your day.
Maybe you should learn about the author before you make such assumptions. I find it hilarious you think he should take statistics courses when he teaches data visualization workshops to places like NASA, IRS, and the UN.
I'm done with this thread. Such a joke.
Just because you’re high profile in the data viz industry doesn’t mean you should be commenting on statistics especially with such a clear misunderstanding going on.
Some of us are definitely more qualified to speak on these matters and we still don’t think we’re qualified to teach it.
He isn’t dogmatic.
He makes reasonable arguments.
I’m confused by these hopelessly uncharitable readings of the article.
"Other reviewers suggested that the conclusion [of this article] should be that box plots are a useful chart type, but only for statistically savvy audiences. Again, I’m going a step further, suggesting that even those audiences would be better served by other chart types in virtually all situations."
A box plot is a data compression technique for compression by hand. There are now better automated techniques that both preserve data quality and visual quality better.
The author is approaching this as a human problem. Plots are not made for machines, they are for people to read, and the author specifically wants as many people can read and parse plots easily as possible. As lamentable as math education might be, we have to work with what we have, and I do think it is a reasonable goal. I agree with the author that it should not be necessary to know what quartiles are in order to see how spread out a distribution is.
Well that explains the entire data visualisation and dashboard consultancy nicely.
How does anyone rationalise the information they have if they don’t make an effort to understand it. Or how can they even select a visualisation method or comparison method. We are truly fucked!
this is precisely why i don't bother with capitalization in my sentences.
in fact even punctuation isnt necessary i dont see why i should dumb down my explanations for people who arent going to make an effort to understand them
actuallyevenspacesaresimplyredundantandasufficientlysmartreadershouldjustunderstandmymeaningwithoutmeneedingtodelineatemywordswhataretheyachildifthiswasgoodenoughfortheancientromansthenitsgoodenoughforme
hckvnvwlsrrdndntndfnynsysthrwsthnmycnclsnsthtthrbrnsrnsffcntlylrgtcmprhndmygns
You can’t control what other people do. You can try to meet them where they are, or hope they’ll catch up with you. Hopefully it’s not your problem if they fail to.
There's generalizations and 'specific situations' which the author considers worthy of some plots, and other specific situations that the author doesn't consider worthy of other plots. At best, don't use box plots if your distributions do not have a single mode and may likely be misinterpreted is my takeaway. Here's a rant against violin plots by my fave physicist ranter[0] (not Sabine), so maybe never use them.
IMO a good takeaway might be to always use a plot that fairly represents the underlying distribution.
https://www.rhoworld.com/i-swarm-you-swarm-we-all-swarm-for-...
Selection of the 'proper' bandwidth is a classic bias-variance tradeoff problem.
* There's another intermediary concept (kernel density estimation) between the audience and the data
* They're still likely to misrepresent tight groupings and discontinuities, which will be smoothed out
These are ok but it’s hard to differentiate the density of points when they’re randomly offset. Try a swarm plot (seaborn) / bee swarm plot (R).
It’s the same concept but the points are strategically placed across the x axis to show the width of the distribution at each point. It generally looks much cleaner.
This is why it's good to have a really competent visual designer around. Their sole purpose is visual communication, and that very much includes dealing with the subconscious connotations and unintended messages hidden within data visualizations. Yes, you've probably encountered designers that would not be good at that, you imagine. You've also probably encountered developers that would not be good at the sort of data munging that scientists, et al do; that doesn't mean developers, generally, aren't best equipped to handle the related coding problems.
It being a staple in statistics is also not a good argument. The information conveyed through box plots is used in lots of fields with different education backgrounds. If a visualization, which in itself is a human simplification of data, is hard to understand, it will be misunderstood by some. This means these people will not be able to advance their field of research as well as with better visualization methodologies.
The author has found that compared to other types of plots, people struggle to learn how to intepret box plots.
The author proposes some alternatives that they believe to be easier for people to interpret:
- Strip plots (for few data points)
- Jittered strip plots (for more data points)
- Distribution heatmap (for even more data points)
----
This aligns with my experience of trying to convey information to non-technical or moderately technical people; box plots are a struggle for them. To me it does seem like the proposed alternatives would be more accessible.
Sure, we could try to better educate people about box plots, (as the author has done professionally); or we could consider using something that requires less effort for people to comprehend.
And still half the comments are like "But I know better!"... yeah, I'd wager most here don't.
You argue in other comments that it's just an education problem, but box plots are used with people who don't have this exact education you mention, and the article explains that a drawback of box plots is exactly that it isn't intuitive and takes several minutes of explanations.
In other words, the article says "I've stopped using this because they require education", and your retort is "Don't stop using these, you just need to educate people".
I wish we could educate everyone in the ways data can be misrepresented - scale, non 0 axis starting, omitting categories, combining groups, colours, point sizes not representative of data - and they can all be levelled at other graph types, singling out box plots for hiding is no different, but IMHO not justification for not using them with the right audience.
[0] https://davidbaranger.com/2018/03/05/showing-your-data-scatt...
As well as using line charts on the average, I've used a box plot (with the edges of the box being the mean +/- 1 standard deviation) to get an idea of whether a given change is significant or not. I.e. if the boxes are close together I will ignore a change I've made, only committing changes that provide a significant jump in performance. The box plot is a useful way of visualizing that.
They can help with seeing highly variable performance (long box) from consistent performance (narrow box).
I can see this in the data (mean, standard deviation) but having it represented visually can help -- especially looking at the data over several iterations, or when looking for patterns from changing a variable (like the number of items in the data being processed).
I've also used linear regression calculations when data has looked linear or quadratic to check/confirm that assumption. -- You can overlay that on top of the data by computing the values for each value of n along side the actual data average and then including the average and calculated values in a line chart.
My fundamental concern with box plots is that no one has ever shown me a single scenario in which a given insight was clearer in a box plot than it would be in a simpler chart type (i.e., strip plot, distribution heatmap, or stacked histograms). If someone can show me even a hand-crafted, cherry-picked scenario with the same data shown as a (well-designed) box plot AND a strip plot, distribution heatmap and stacked histograms, and in which a potentially useful insight is clearer in the box plot than in the other chart types, I’ll happily change my opinion. I’m still waiting for someone to show me such a scenario, though.
In the meantime, I’m not sure why one would use box plots when simpler chart types are available that say the same thing about the data or, in many cases, say more about the data (show gaps, multi-modal distributions, etc.). Even if the audience is very used to reading box plots, they’ll still find strip plots, distribution heatmaps and stacked histograms to be simpler to read (and will actually see gaps, clusters, etc.)
How do I know that other distribution chart types are simpler to read than box plots? Because I’ve taught these chart types to literally thousands of people of all skill levels all over the world. Quartiles are just inherently less intuitive than bins or, in the case of strip plots, no delimiters to understand at all.
Like I said, if someone can show me a scenario like the one that I described above, though, I’ll happily change my mind…
Before people jump all over me, I should clarify what I mean by a “potentially useful insight.” For example, “showing the interquartile range” is not an “insight” in this context, it’s an “observation” because it doesn’t point to any kind of action or conclusion, in and of itself. A potentially useful insight would be something like, “The employee salaries in Company A are generally higher than those in Company B.” or “Most people make close to $80K in Company A, but the salaries are much more spread out in Company B.” Basically, an “insight” in this context is a piece of information that would point directly to some kind of action or conclusion.
https://en.wikipedia.org/wiki/Sina_plot
https://cran.r-project.org/web/packages/sinaplot/vignettes/S...
Adding on a representation of mean in a different style (like a black bar) can be helpful. So can a boxplot-style indication of variance, in some cases.
For example, though the final example in the reference there is graphically "only" shading the "outer band" darker than the inner alpha-blended region, this seems important statistically/visualization-wise since the unknown true parent distribution/ensemble samples are, well, sampled from need only be any monotonic curve within the whole region.. (not even differentiable if mixed discrete-continuous values may happen).
On gnu-R:
install.packages('ggplot2')
?ggplot2::geom_violin
- Bar
- Scatter
- Line
- Histogram
You can tell 90% of your stories with these plots. (If you pay attention to professional viz groups, Economist, NY Times, etc, they use these.)
Don't waste your time with other plots unless you have mastered these. When you master these, you will realize you don't need other charts.
The above are between the reasons I prefer remote meeting where data are to be shown instead of in person: anyone attending should have a computer ready to use and IF data are shared and ready usable I can live tweaks a plot ad reason on it while I listen end eventually pose relevant questions shown at my own turn something. Surely not all presentations are meant to be interactive session, but being able to interact even in async form reading a journal article, playing with the data and eventually drop a mail to the author is a nice thing, typically uselessly hard today where in tech term it can be extremely simple.
That's another reason I have presentation software/office automation one instead of plain org-mode, Jupyter, R Studio etc because change things it's hard while it should be easy. Org-mode is excellent to present but not really interactive, I have to regenerate plots to see changes or push data to external software, Jupyter is not really meant to present, R Studio offer nice LaTeX integration and tabular view but do not offer nice means to present, though they are still FAR better then presentation software and even if have some safety aspects to be taken into account I prefer countless of time receiving an active document (org-mode, jupyter notebook etc) instead of a pdf or even worse some office formats.
Just plotting points will lead to saturation in high density areas that depends on point size and opacity.
Making bin color proportional to point density will require normalization to make the plot readable in many cases.
While I like these plots too in certain situations, I would argue they're actually less elegant than the boxplots for those reasons.
And come on, boxplots aren't that hard to explain to someone who already is used to working with percentiles.
But that's okay, I don't mind explaining it and then the graph is easier to interpret imo.
Of course you do: they're in the whiskers; half in each whisker.
That's the entire point of the picture, BTW.
That's like the entire point of the post: they're hard to teach to others (they're unintuitive) and there are better (more intuitive) alternatives.
I dunno if I agree, but it's ironic that this thread started with a poster complaining about the author's bad intuition, while apparently managing to not have a good grasp of box plots themselves.
Also, that IS his mistake, it's literally the first thing in the post. And this stuff isn't hard or hard to teach _at all_ has long as you're at least 5.
> You don't know where the rest are.
This is wrong, period. And the fact it's wrong is pretty much the entire point of the article.
Are bdjsiqoocwk and lkdfjlkdfjlg the same poster?
Please don't pick a needless fight.
Apart from that quibble, it's a point very well taken.
Isn't that a bit like banning cars because some people can't drive?
Some diagrams are simply not for mass consumption and this is one, particularly because it is designed to illustrate an interpretation of ranges instead of the direct/linear representation of the raw data.
Of course I'd illustrate this fact as a Venn diagram comparing "box diagram" Vs "people" (intersection those who understand it) but I'm afraid the universal set may be mistaken as "those people who don't have eyes" rather than literally everything else.
Perhaps we should stop using that too, since it's non obvious what the universal set is.
All diagrams have some ambiguity and can be misinterpreted, sometimes it's deliberate (e.g. bar chart vertical axis not starting at 0 or scale not being linear) and that's why there's the saying "There's lies damn, lies, and statistics." That doesn't mean some diagrams are not useful, just that it's not suitable for some audiences who may misinterpret the data.
We have better technology nowadays, including for plotting, so why not ditch the old?
The author of the blog post has some good arguments. From your post, I cannot distill an argument as to why you would prefer specifically a box plot over a strip plot.
But that's not the purpose of a box diagram and the article even did a side by side comparison showing an apples and oranges comparison of 2 total different representations of the data.
Those diagrams were never meant to represent the data in the same way.
The article simply could have shown a better way of illustrating the data, rather than implying box diagrams are incorrect, which they aren't, any more than choosing a bad graph or axis is (CF. parent comment)
Sometimes you may want to highlight some core representation of data without the distraction of outliers (yes that does mean some people will use it for deliberate misrepresentation). But in this regard it's useful, as is on bar graphs not starting the vertical at 0 (because you want to illustrate relate difference not absolute amounts).
If a medium of communication is misunderstood and found to be misleading to your audience, it doesn't really matter whether it's an education problem or not. It ceases to be a good communication medium.
The entire purpose of data viz as the author discussed is to convey ideas to other people. The author argues that people tend to misunderstand this specific chart type. It is valid, then, to dismiss the visualisation as bad for public communication.
Unfortunately, the technical merits of these things don't matter if most people don't understand them.
There are the better graphs the author mentioned for general purpose use, but the graph itself isn't at fault any more than using a bar chart with a poor scale (e.g omit 0-20) to do the same hiding.
There is a lot of implicit (e.g. traffic signals) and explicit (e.g. indicators and horns) inter-driver communication that is at the heart of most crashes.
But the real education deficit shown here is psychology education. Humans are bad at doing some calculations inherently. They are not able to properly asses pie charts and easily confused by numbers with a lot of digits. Even before these studies were done people were able to come up with visualizations that were better suited for human understanding.
People chose to use box plots because the visualization was better to understand by people than the numerical representation of the same information. Luckily there are now even better tools to represent the same numerical data in a way that is even better to understand.
So, if you are truly educated properly you don't use visualization.
A problem like this one that he mentions, "People associate longer shapes with greater quantity", is not something you can fix by teaching. Even if you know intellectually that the association is, in this case, wrong, you can't free yourself from the association. It's hardwired into the brain.
People who work with this sort of diagram a lot will eventually build up context-specific associations that work better, overriding that instinct, to the point where it feels seamless. But even if it feels seamless and easy, the dissonance is still there, and may lower your comprehension speed and slightly impair your judgment.
As a statistics expert, you are never going to notice that, because your baseline comprehension speed and judgment on the subject is so good, that this very minor impairment is lost in the noise. So you may not be a good judge of the usability qualities of the diagram type.