Violin plots + outlier dots + additional markers are more helpful. Sometimes the CDF is also more useful than the PDF, e.g. for latencies.
It's basically just my boss making suggestions and me implementing, so the results are probably less than optimal for this kind of thing.
[0] is just one of the first results from a search that makes this case.
[0] https://stanfordreview.org/calculus-is-overrated-why-we-shou...
I like swarm plots for this kind of task.
https://seaborn.pydata.org/generated/seaborn.swarmplot.html
Edit: well, I suppose I should clarify the comment on violin plots is implementation dependent and biased by my personal preferences for visualization libraries
Showing the number of samples is good though. The combined version seems useful.
Consider inner="stick" from [1] instead. It probably comes closer to what you're looking for.
[1] https://seaborn.pydata.org/generated/seaborn.violinplot.html
Like all data visualization techniques, which plot type is better depends on the context.
https://wellcomeopenresearch.org/articles/4-63/v1
They can visualize common statistics like box-plots, the distribution shape like violin plots, as well as the raw data!
The important thing here is understand the underlying system and identify what it's modes of failure are. If the likely mode of failure was just to fail to serve some requests at all then the new approach in this article would be entirely useless since the latencies would be the wrong metric completely.
In this case they identified one mode of failure is for a small % of requests to deteriorate almost exponentially and so watching the 99th percentile is useful - this is very common in networking. That's great for this system but the overall message is to find the metric that actually reflects the ways your system will fail- and you need to do that in a practical way that may not necessarily involve recording and storing all the data.
My manager's dictum is histogram and time series to start, which falls under the same auspices.
I think percentiles over time, given consistent sample sizes per bucket, can be good for understanding change over time.
In general I think of histograms for these sort of metrics providing higher fidelity of value, where as stats over time can provide higher fidelity of time (for roughly the same storage).
Animating the distribution over time can be pretty neat, but not common feature in my space (but can be done with a little R code).
I imagine 3 axis (3d) histograms over time can make some neat visualizations but have never experimented with that, and would probably want to slice or rotate the plot. More common are heatmaps where the darkness of the color is your third axis. Kind of like viewing the 3d plot of histograms over time from above where the cubes are partly transparent and get darker the deeper they are.