When assessing probabilistic models the plots should be showing the mean a̶n̶d̶ ̶s̶t̶d̶e̶v̶ of many monte carlo simulations not just one line per model and claiming "look this model is more gooder!"
common mistake people make
P(|X-\mu| > k \sigma) < 1/k^2.
So, while for a normal RV, 5% of observations lie outside +/- 1.96 std.devs, for arbitrary RV (with finite variance) at most 25% of observations lie outside +/- 2 std.devs.