Forecasts need to have error bars
andrewpwheeler.com
andrewpwheeler.com
(2) it's not clear what the "error bars" should mean. One is a confidence interval[1] (e.g. model gives 95% chance the output will be within these bounds). Another is a standard deviation (i.e. you are pretty much predicting the squared difference between your own point forecast and the outcome).
[1] acknowledged: not the correct term
Does that mean you should be more surprised if your predictions are wrong? It depends. You've only reduced model error, but this is the classic precision versus accuracy problem. You can very precisely estimate the wrong number. Does the model really reflect the underlying data generating process? Are the inputs you're giving reliable measurements? If both answers are yes, then your more precise model should be converging toward a better prediction of the true value, but if not, you're only getting a better prediction of the wrong thing.
We can ask these questions with this very example. Clearly, ARIMA is not a causally realistic model. Criminals don't look at last year's crime rates and decide whether to commit a crime based on that. The assumption is that, whatever actually does cause crime, it tends to happen at fairly similar levels year to year, that is, 2020 should be more different than 2010 than it is different from 2019. We may not know what the causually relevant factors really are or we may not be able to measure them, but we at least assume they follow that kind of rule. This sounds plausible to me, but is it true? We can backtest by making predictions of past years and seeing how close they are to the measured value, but the possibility of this even working depends upon the answer to the second question.
So then the second question. Is the national violent crime data actually reliable? I don't know the answer to that, but it certainly isn't perfect. There is a real crime rate for every crime, but it isn't exactly the reported number. Recording and reporting standards vary from jurisdiction to jurisdiction. Many categories of crime go undereported, the extent of which can change over time. Changes may reflect different policing emphasis as much as or more than changes in the underlying true rate. I believe the way the FBI even collects and categorizes data has changed in the past, so I'm not sure a measurement from 1960 can be meaningfully compared to a measurement from 2020.
Ultimately, "how surprised you should be when you are wrong" needs to take all of these sources of error into account, not just the model's coefficient uncertainty.
When trying to detect cheating in online games you don’t need to predict exact performance, but you want to decent anomalies quickly. Detecting cereal killers, gang wars, etc isn’t about nailing the number of murders on a given day but patterns within those cases etc.
You only really need to take those sources of error into account if you want an absolute measure of error, which as you explain, seems pretty impossible.
An error for weather only needs to be relative -- for example, if the error for rain today is higher than yesterday, it's not important the exact number higher that it is -- only that it's higher. (Not that I know if this is possible.)
It's like how you can't describe how biased a certain news source is or how to read a Yelp or Rotten Tomatoes review -- you just have to read it often enough to get an intuitive sense that a 4.1 star Yelp-rated restaurant with 800 reviews is probably good while a 4.6 star restaurant with 5 reviews is quite possibly terrible.
"You should be willing to take either side of the bet that confidence interval implies." (paraphrasing; he says it better).
For a concrete example, with a 95% confidence interval, you should be as willing to accept the 19:1 odds that the true value is outside the interval as you are the 1:19 odds that the true value is inside the interval.
Aside from being generally correct, this approach is immediately actionable by making the meaning more visceral in discussions of uncertainty. Done right, it pushes you to assign uncertainties that are neither too conservative nor too optimistic.
If the notion of letting your reader take either side of the bet makes your stomach a little queasy, you're on the right track. The feeling will subside when you're pretty sure you got the errorbar right and your reasoning is documented and defensible.
Edit for OP's explicit question: One standard-deviation errorbars are 68% confidence intervals. Two standard deviations are 95% confidence intervals. (assuming you're a frequentist, of course)
[1] https://www.nobelprize.org/prizes/physics/1997/phillips/fact...
Also assuming normal distribution, I think?
> If the notion of letting your reader take either side of the bet makes your stomach a little queasy, you're on the right track. The feeling will subside when you're pretty sure you got the errorbar right and your reasoning is documented and defensible.
> For a concrete example, with a 95% confidence interval, you should be as willing to accept the 19:1 odds that the true value is outside the interval as you are the 1:19 odds that the true value is inside the interval.
I would like to build some edge into my bets. If a reader takes both sides of your example, they would be come out exactly even.
But since readers are not forced to take any side at all, they will only take the bet if one of the sides has an advantage.
So I would like to be able to say, 20:1 payout the true value is inside the error bar, and 1:10 payout it's outside the error bar (or something like that).
The tighter the spread I am willing to quote, the more confident I am that I got the error estimates right. (I'm not sure how you translate these spreads back into the language of statistics.)
95% is 95% regardless of the distribution.
> I would like to build some edge into my bets. If a reader takes both sides of your example, they would be come out exactly even.
You can imagine yourself being equally unhappy to take either side of the bet, if that's easier than imagining yourself being happy to take either side.
It is for me, which is probably something to bring up in therapy.
I also think that framing things as bets brings in all the cultural baggage around gambling and so it isn't always helpful. I'm not sure what a better framing is though.
Standard deviation away from the mean don't correspond to the same percentiles for all distributions, or do they?
If you want to be (almost) independent of distribution, you need Chebyshev's inequality. But that one is far weaker.
> Its practical usage is similar to the 68–95–99.7 rule, which applies only to normal distributions. Chebyshev's inequality is more general, stating that a minimum of just 75% of values must lie within two standard deviations of the mean and 88.89% within three standard deviations for a broad range of different probability distributions.[1][2]
https://en.wikipedia.org/wiki/Chebyshev%27s_inequality
> I also think that framing things as bets brings in all the cultural baggage around gambling and so it isn't always helpful. I'm not sure what a better framing is though.
Underwriting insurance without going bankrupt, perhaps?
If they had said a two standard deviation interval then you would have needed to know the distribution, but they said 95% which gives you all the information you need to make the bet.
> Edit for OP's explicit question: One standard-deviation errorbars are 68% confidence intervals. Two standard deviations are 95% confidence intervals. (assuming you're a frequentist, of course)
You are right about the first part of the original comment.
(Of course, that's not as good as betting money.)
Many of them eg work in universities and some even have tenure. There's not much skin in the game between any forecasts they might make and their academic prospects.
Economists working for companies often have to help them understand micro and macro-economics. Eg (some of) Google's economists help them design the ad auctions. It's relatively easy to figure out for Google how well those ad auctions work. So they certain have skin in the game. But: what motivation to mislead do those economists have?
How do you know that? Whenever I interact with economists, mostly online via blogs but also sometimes via email, they always seem painfully aware of the shortcomings of their models, and don't seem to confuse them with reality.
Perhaps you have studied a different sub-population of economists than the ones I have anecdotal experience with?
Bought into is not the same as believing.
Why do physicists ignore friction whenever possible?
In general, for any task, you take the simplest model that represents the aspects of reality that you care about. But you stay aware of the limits. That's true in physics or engineering just as much as in economics.
That's why NASA uses Newtonian mechanics for all their rocket science needs, even though they have heard of General Relativity.
That's why people keep using models known to have limits.
> [...] sage voices echo the standard dogma [...]
You do know that most of published economics is about the limits of the 'standard dogma'? That's what gets you published. I often wish people would pay more attention to the orthodox basics, but confirming well-known rules isn't interesting enough for the journals.
So if eg you can do some data digging and analysis that can show that maybe under this very specific circumstances restriction on free trade might perhaps increase national wealth, that can get you published. But the observation that most of the time free trade, even if the other guy has tariffs, is the optimal policy, is too boring to get published.
Compare also crap like 'Capital in the Twenty-First Century' that catapults its author to stardom with its comparatively boring refutation by orthodox economists that no one cares about.
> [...] dragged into the doldrums by policy posed by useless models.
Most orthodox economics is pretty unanimous about basic policies: for free trade, against occupational licensing, for free migration, for free movement of capital, for simple taxes without loopholes, against messing with the currency, against corruption, against subsidies, for taxes instead of bans (eg on drugs, or emissions, or guns), against price floors or ceilings or other price controls, etc.
Many doldrums happen when policy ignores or contradicts these basic ideas. Alas, economics 101 is not popular with the electorate almost anywhere.
For example you say a basic policy is a tax on guns instead of a ban. First of all I dispute that is even orthodox economics. Second there is some strong evidence that gun bans reduce violence.
Free migration is another one. It is an insanely complicated issue in the real world. No country has 100% free migration or they wouldn’t be a country. There are all kinds of very complex rules and effects of these rules. And it is not clear that “free migration” is “good”. (I am sure the native americans probably didn’t like the free migration)
First, I apologize for using guns as an example. That's a needlessly divisive topic. The general principle of 'taxes instead of bans' is rather orthodox. You see that more often applied to the example of drugs or emissions.
Second, what evidence do you have for gun bans reducing violence? And reduce violence compared to what baseline?
I am very willing to believe that if you compare a free-for-all with a ban on guns, that the latter will see less violence. (I haven't looked into the evidence. Results might differ depending on details and on when and where you do that, and who gets exceptions to the bans. Eg police and military presumably are still allowed guns? Hunters probably as well? Etc. It's not so important.)
My point is that in terms of violent crime avoided, a situation where each gun and each bullet comes with a million dollar tax would be statistically indistinguishable from a ban.
And in practice, a less severe tax would probably be enough to achieve those goals whilst still preserving access to guns for those who prefer it that way.
> No country has 100% free migration or they wouldn’t be a country.
What kind of definition of 'country' are you using here that breaks down in this way? (And what do you mean by '100%'? How nitpicky do you want to be?)
A history lesson from Wikipedia https://en.wikipedia.org/wiki/Passport
> A rapid expansion of railway infrastructure and wealth in Europe beginning in the mid-nineteenth century led to large increases in the volume of international travel and a consequent unique dilution of the passport system for approximately thirty years prior to World War I. The speed of trains, as well as the number of passengers that crossed multiple borders, made enforcement of passport laws difficult. The general reaction was the relaxation of passport requirements.[18] In the later part of the nineteenth century and up to World War I, passports were not required, on the whole, for travel within Europe, and crossing a border was a relatively straightforward procedure. Consequently, comparatively few people held passports.
> During World War I, European governments introduced border passport requirements for security reasons, and to control the emigration of people with useful skills. These controls remained in place after the war, becoming a standard, though controversial, procedure. British tourists of the 1920s complained, especially about attached photographs and physical descriptions, which they considered led to a "nasty dehumanisation".[19] The British Nationality and Status of Aliens Act was passed in 1914, clearly defining the notions of citizenship and creating a booklet form of the passport.
Btw, Switzerland as a country does not restrict immigration. That's left to the Kantone (which are sort-of the equivalent to American states). Yet, you'd be hard pressed to argue that Switzerland is not a country. If memory serves right, the US used to have similar arrangements in their past?
It's interesting that if you oblige your models to fit a set of policy positions then they return that set of policy positions and are pretty useless in general. A cynic might say that's by design.
Orthodox macroeconomic modelling is laughably naive and mathematically wrong before even getting to the basic issues of failure to validate. Let's not compare it to disciplines where validation is the entire point.
Your rhetoric clearly shows you don't want to think too critically about this so I'll sign off now.
Input: temperature is measured at -1° C.
Prediction: water will freeze.
Actual: water didn't freeze.
Actual temperature: 2° C.
The model isn't broken, it gives an incorrect result because of input error.
The model itself can perfectly describe the physics, but it only knows what you can give it. This may be limited by measurement uncertainty of your equipment, etc, but it is separate from the model itself.
In this area, "the model" is typically considered as the input parameter to quantity of interest map itself. It's not the full problem from gathering data to prediction.
Model error would be things like failing to capture the physics (due to approximations, compute limits, etc), intrinsic aleatoric uncertainty in the freezing process itself, etc.
Making this distinction helps talk about where the uncertainty comes from, how it can be mitigated, and how to use higher level models and resampling to understand its impact across the full problem.
I say in theory because, in my experience in the tech industry, with the usual exceptions, uncertainty intervals, for example on a graph, are interpreted by those making decisions as aesthetic components of the graph ("the gray bands look good here") and not as anything even marginally related to a prediction.
I have more doubts when it comes to actions taken when considering properly estimated predictive intervals. Even I, who have a good knowledge of statistical modeling, after hearing "the median survival time for this disease is 5 years," do not stop to think that the median is calculated/estimated on an empirical distribution, so there are people who presumably die after 2 years, others after 8. Well, that depends on the variance.
But if I am so strongly drawn to a central estimate, is there any chance for others not so used to thinking about distributions?
True. But a conditional quantile is much harder to accurately estimate from data than a conditional expectation (particularly if you are talking about extreme quantiles).
If you don't mind typing it out, what do you mean formally here?
If you think about linear regression, it makes sense, given the assumptions of linear regression, that confidence interval E[x|y] is narrower around the mean of x and y.
If I had to choose between the two, confidence intervals in a forecasting context are less useful in the context of decision-making, while prediction intervals are, in my opinion, always needed.
And, from what I understand, this is what is happening in this article.
The person is providing an uncertainty interval for their mean estimator and not for future observations (i.e., the error bars reflect the uncertainty of the mean estimator, not the uncertainty over observations).
Like you said: before adding error bars, it probably makes sense to think a bit about what type of uncertainty those error bars are supposed to represent.
And it's very different from what I expected, and it doesn't make a lot of sense to me. I guess if statisticians already believe your model, then they want to see the error bars on the model. But I would expect if someone gives me a forecast with "error bars", those would relate to how accurate they think the forecast would be.
Though probabilistic forecasting methods may have a Bayesian approach, it is Monte Carlo sampling that helps generate the confidence intervals.
Feel free to correct me if I am wrong! Thanks :)
What you probably want is the standard error, because you are not interested in how much your data differ from each other but in how much your data differ from the true population.
ARIMA is based on a few assumptions. One, there exists some "true" mean value for the parameter you're trying to estimate, in this case violent crime rate. Two, the value you measure in any given period will be this true mean plus some random error term. Three, the value you measure in successive periods will regress back toward the mean. The "true mean" and error terms are both random variables, not a single value but a distribution of values, and when you add them up to get the predicted measurement for future periods, that is also a random variable with a distribution of values, and it has a standard error and confidence intervals and these are exactly what the article is saying should be included in any graphical report of the model output.
This is a characteristic of the model. What you're asking for, "how wrong do you think the model is," is a reasonable thing to ask for, but different and much harder to quantify.
> What you're asking for, "how wrong do you think the model is," is a reasonable thing to ask for, but different and much harder to quantify.
This definitely seems to me to be what the original author is motivating: forecasts should have "error bars" in the sense that they should depict how wrong they might be. In other words, when the author writes:
> Point forecasts will always be wrong – a more reasonable approach is to provide the prediction intervals for the forecasts. Showing error intervals around the forecasts will show how Richard interpreting minor trends is likely to be misleading.
The second sentence does not sound like a good solution to the problem in the first sentence.
Just to clarify... this is Python code, not R.
The intuition would be that some% of your forecasts are between the bounds of the credible interval.
Some thought it clouded the issue. For example, when a new treatment caused a 1% "improvement", but the confidence interval extended from -10% to 10%, it was clear that the experiment didn't tell us how that metric was affected. This makes the decision feel more arbitrary. But that is exactly the point - the decision is arbitrary in that case, and the confidence interval tells us that, allowing us to focus on other trade-offs involved. If the confidence interval is 0.9% to 1.1%, we know that we can be much more confident in the effect.
A big problem with this is that meaningful error bars can be extremely difficult to come by in some cases. For example, imagine having something like that for every prediction made by an ML model. I would love to have that, but I'm not aware of any reasonable way to achieve it for most types of models. The same goes for online experiments where a complicated experiment design is required because there isn't a way to do random allocation that results in sufficiently independent cohorts.
On a similar note, regularly look at histograms (i.e., statistical distributions) for all important metrics. In one case, we were having speed issues in calls to a large web service. Many calls were completing in < 50 ms, but too many were tripping our 500 ms timeout. At the same time, we had noticed the emergence of two clear peaks in the speed histogram (i.e., it was a multimodal distribution). That caused us to dig a bit deeper and see that the two peaks represented logged-out and logged-in users. That knowledge allowed us to ignore wide swaths of code and spot the speed issues in some recently pushed personalization code that we might not have suspected otherwise.
This is something I've started noticing more and more with experience: people really hate arbitrary decisions.
People go to surprising lengths to add legitimacy to arbitrary decisions. Sometimes it takes the shape of statistical models that produce noise that is then paraded as signal. Often it comes from pseudo-experts who don't really have the methods and feedback loops to know what they are doing but they have a socially cultivated air of expertise so they can lend decisions legitimacy. (They used to be called witch-doctors, priests or astrologers, now they are management consultants and macroeconomists.)
Me? I prefer to be explicit about what's going on and literally toss a coin. That is not the strategy to get big piles of shiny rocks though.
This is extremely common and one of the core ideas of statistical process control[1].
Sometimes you have just the one process generating values that are sort of similarly distributed. That's a nice situation because it lets you use all sorts of statistical tools for planning, inferences, etc.
Then frequently what you have is really two or more interleaved processes masquerading as one. These distributions generate values that within each are sort of similarly distributed, but any analysis you do on the aggregate is going to be confused. Knowing the major components of the pretend-single process you're looking at puts you ahead of your competition -- always.
[1]: https://two-wrongs.com/statistical-process-control-a-practit...
They are pretty much a one sided distribution power law. Deadlines almost never come in early and, when they do, it's rarely by much. On the other hand, deadlines can come in late by wild amounts.
Generating confidence intervals on that is really hard.
A date estimation with no error bars cannot be proven wrong. But! If you say "there's a 50 % chance it's done before this date" then you can look back at your 20 most recent such estimations and around 10 of them better have been on time. Otherwise your estimations are not calibrated. But at least then you know, right? Which you wouldn't without the error bars.
I regularly see two kinds of deadlines. "Planning deadlines" describe an estimate when something will be done. "Signalling deadlines" signal priorities and motivation to employees or clients. Sometimes both exist in parallel for the same task and there is a subset of people who know both.
I always demand error bars.
Another, separate, issue that is often neglected is the idea of calibrated model outputs, but that's its own rabbit hole.
For instance, if you look at https://blog.tensorflow.org/2019/03/regression-with-probabil... until the case 4 it's easy to follow and digest, but if you look at the _Tabula rasa_ section I am pretty sure that such content isn't understandable by many. Where you get stuck because the ideas become too complex depends on your math skills.
I guess my point is, there is no silver bullet. Adding defensible uncertainty is complicated and problem specific, and comes with downsides (often steep).
Sure, you'll ideally want a calibrated estimator/superforecaster to do it, but they exist and they aren't that rare. Any decently sized organisation is bound to have at least one. They just need to care about finding them.
I do think having people in the loop is a very important aspect, however, and can provide an important subjective complement to the more mathematically formulated idea of uncertainty. I don't care if the model I'm using provides the most iron clad and rigorous uncertainties ever, I'm still going to spot check it and play with it before I consider it reliable.
By having their forecasts continuously evaluated against outcomes. If someone can show me they have a track record of producing calibrated error bars on a wide variety of forecasts, I trust them to slap error bars on anything.
> even if a human is able to slap an uncertainty on a prediction [...] that doesn't mean it's representing the uncertainty of what the model based its decision on.
This sounds like it's approaching some sort of model mysticism. Models don't make forecasts, humans do. Humans can use models to inform their opinion, but in the end, the forecast is made by a human. The human only needs to put error bars on their own forecast, not on the internal workings of the model.
By forecasts I only mean output of a model, I've been wrapped up in time series methods where that's the usual term for model outputs. Assigning confidence to the conclusions drawn by an analyst using some model as a tool is a different task that may or may not roll up formal model output uncertainties and usually involves a lot of subjectivity. This is an important thing too, but is downstream.
Uncertainty is inherently tied to a specific model, since it characterizes how the model propagates uncertainty of inputs and its own fit/structure/assumptions onto its outputs. If you aren't building uncertainties contingent on the characteristics of a specific model then it isn't an uncertainty. But there's no mysticism about models possibly being unintuitive, most of the popular model forms nowadays are mystery black boxes. Some function fit to a specific dataset until it finds a local minimum in a loss function that happens to do a good job (simplifying). There's plenty of work that shows ML models often exploit features and correlations that are highly unintuitive to a human or are just plain spurious.
When I used to publish stats- and math-heavy papers in the biological sciences, very rarely the reviewers--and I used to publish in intermediate and up journals--were paying any attention to the quality of the predictions, beyond a casual look at the R2 or R2-equivalents and mean absolute errors.
The goal of a business is usually to generate revenue so I'm confused about your question.
any measurement that you make without any knowledge
of the uncertainty is meaningless
https://youtu.be/6htJHmPq0OsYou could say that forecasts are measurements you make about the future.
"Being able to quantify uncertainty, and incorporate it into models, is what makes science quantitative, rather than qualitative. " - Lawrence M. Krauss
It turns out the ECMWF does do an ensamble model where they run 51 concurrent models, presumably with slightly different initial conditions, or they vary the model parameters within some envelope. From these 51 models you can get a decent confidence interval.
But this is a lower resolution model, run less frequently. I assume they don't do this with their "HRES" model (which has twice the spacial resolution) in an ensemble because, well, it's really expensive.
[1]: https://en.wikipedia.org/wiki/Integrated_Forecast_System#Var...
ECMWF recently upgraded their ensemble to run at the same resolution as the HRES. The HRES is basically the ensemble control member at this point [1]
[1] https://www.ecmwf.int/en/about/media-centre/news/2023/model-...
The most fascinating thing about the concept of error in models, including in ensembles, is you can only calculate and propagate error for contributors that you can quantify. There are many unquantifiable sources of error. Imagine a physical process that you are unaware of that propagates as a bias, for example ice nucleation via aerosols. Perhaps you don't even model aerosols. How do you account for error here? What does error even mean?
Ensembles only show you intramodel variability. Which is like error, sort of, but only really represents a combination of "real" variability in initial conditions and how that propagates through your physics/parameterizations.
"models" the HN commentators make for their businesses surely have parallel concepts, but I don't see anyone talking about them. Only discussion about the errors you know when the ugliest errors are the ones that no one knows.
https://content.meteoblue.com/en/research-education/specific...
They also have one for precipitation type distribution: https://charts.ecmwf.int/products/opencharts_ptype_meteogram...
https://weather.gc.ca/city/pages/on-143_metric_e.html
I really like that discussion type forecast.
Up until that point, error bars increase. At least to me, there's a big difference between "1 mm rain guaranteed" and "90 % chance of no rain but 10 % chance of 10 mm rain" but both have the same average.
The result is that I will be told ~50% chance of precipitation, but the places I care about might well be in the essentially 0% or essentially 100% parts.
The single forecast for a decent size city problem impacts other parts like forecasted highs/low. Even without a front, in some cities that have incorporated much of their suburbs, it is not uncommon for the eastmost part and westmost part to differ by 5-6 degrees, and that is ignoring the inherent temperature differences found in the downtown areas.
I've been seeing this question come up a lot lately. The answer is no, weather forecasting continues to improve. The rate is about 1 day improvement every 10 years so a 5 day forecast today is as good as a 4 day forecast 10 years ago.
It's sloppy science / statistics to not haven error ranges.
> An illusion of predictability in scientific results: Even experts confuse inferential uncertainty and outcome variability
> Traditionally, scientists have placed more emphasis on communicating inferential uncertainty (i.e., the precision of statistical estimates) compared to outcome variability (i.e., the predictability of individual outcomes). Here, we show that this can lead to sizable misperceptions about the implications of scientific results. Specifically, we present three preregistered, randomized experiments where participants saw the same scientific findings visualized as showing only inferential uncertainty, only outcome variability, or both and answered questions about the size and importance of findings they were shown. Our results, composed of responses from medical professionals, professional data scientists, and tenure-track faculty, show that the prevalent form of visualizing only inferential uncertainty can lead to significant overestimates of treatment effects, even among highly trained experts. In contrast, we find that depicting both inferential uncertainty and outcome variability leads to more accurate perceptions of results while appearing to leave other subjective impressions of the results unchanged, on average.
[1] https://www.microsoft.com/en-us/research/publication/an-illu...
"Point forecasts will always be wrong" - true that for continuous data but if you can predict that some stock will go to 2.01x it's value instead of 2x that's still helpful.
https://en.wikipedia.org/wiki/Gaussian_process#Gaussian_proc...
Another famous hypothesis is the phasing out of lead fuel: https://en.wikipedia.org/wiki/Lead%E2%80%93crime_hypothesis
For instance in a business setting, if I say "it'll be done in 10 days +/- 4 days", they'll immediately say "ok so you're saying it'll be done in 14 days tops then".
More effective to sound as unsure as possible, disclaim everything in slippery language, and promise to give updates to your predictions as soon as you realise they've changed (granted this wouldn't work as well for an anonymous reader situation like in this article).
Accounting should do it too in their reporting.
I would love to see a balance sheet with a proper 'certainty range' around the values in there.
I would suggest split conformal first.
Its depressing
Why, I hear you ask? Because, for the kind of system of models I use (detailed stochastic simulations of human behavior), there is no good definition of a confidence interval that can be computed in a reasonable amount of computing time. One can design confidence measures that can be computed without too much overhead, but they can be misleading if you do not have a very good understanding of what they represent and do not represent.
To simplify, the error bars I was able to compute were mostly a measure of precision, but I had no way to assess accuracy, which is what most people assume error bars mean. So showing the error bars would have actually given a false sense of quality, which I did not feel confident to give. So not displaying those measures was actually done as a service to the user.
Now, one might make the argument that if we had no way to assess accuracy, the type of models we used was just rubbish and not much more useful than a wild guess... Which is a much wider topic, and there are good arguments for and against this statement.
It seems like the options are:
- no error bars which mislead everyone
- error bars which confuse some people and accurately inform others
See also: Complaints about poll results in the last few rounds of elections in the US. "The polls said Hillary would win!!!" (no, they didn't).
It's not just error margins, it's an absence of statistics of any sort in secondary school (for a large number of students).
{Confidence interval we won't cook the planet}