Prophet: Automatic Forecasting Procedure
github.com
github.com
They recommend checking out these for cutting-edge time series forecasting:
Another thing that both NeuralProphet and Prophet do extremely wrong by default is uncertainty estimation. The coverage probabilities are way off.
It generated very efficient samplers for particularly weird (and enormous!) hierarchical models I had. Documentation is also great.
It is also worth reading Andrew Gelman's post about Prophet: https://statmodeling.stat.columbia.edu/2017/03/01/facebooks-...
The problem with time series forecasting in general is that they make a lot of assumptions on the shape of your data, and you'll find you're spending a lot of time figuring out mutating your data. For example, they expect that your data comes at a very regular interval. This is fine if it's, say, the data from a weather station. This doesn't work well in clinical settings (imagine a patient admitted into the ER -- there is a burst of data, followed by no data).
That said, there's some interesting stuff out there that I've been experimenting with that seems to be more tolerant of irregular time series and can be quite useful. If you're interested in exchanging ideas, drop me a line (email in my profile).
Does prophet rely on this assumption? For health timeseries data the tool of choice is survival analysis - typically using Cox proportional hazards regression or similar regression tools that are able to handle irregular or censored data.
I've seen some moves towards using fancy bayesian or fancier machine learning stuff for clinical trials but a big issue is that they are very difficult to communicate to their intended audience.
(Thought the actual sampling mechanics and tooling can be much more complex)
Re: "fancier machine learning" -- I've seen different flavors of RNNs & LSTMs have some success in analyzing time series data. I've struggled to get them to work on real-world (i.e., messy) data, but have had some encouraging results with a transformer encoder-only NN.
Also, Prophet was developed by a very small number of individuals at Facebook, it's not something they invested massive resources into.
A GP is an intuitive and expressive way to code time covariance in a model. A famous example is the relative birthdays model, discussed by Gelman et al in Bayesian Data Analysis and here [1].
[1] https://avehtari.github.io/casestudies/Birthdays/birthdays.h...
[disclaimer I'm a maintainer of Hamilton] Otherwise FYI Prophet gels well with https://github.com/DAGWorks-Inc/hamilton for setting up your features and dataset for fitting & prediction[/disclaimer].
Curiously, in Medium-like (ie low effort) publications it's still the recommended way to tackle a forecasting problem. The promise of a model that can solve any time series problem sounds great, but not all that glitters is gold, and as you get more experience you discover that solutions like this usually don't work.
[1] - https://ryxcommar.com/2021/11/06/zillow-prophet-time-series-...
Zillow, Prophet, time series, and prices - https://news.ycombinator.com/item?id=29137200 - Nov 2021 (143 comments)
Is Facebook's “Prophet” the time-series Messiah or just a naughty boy? - https://news.ycombinator.com/item?id=27695574 - July 2021 (78 comments)
I'll disclaim that I'm just a finance dude and not a data scientist or programmer. But the documentation leads me to believe that I am in the target audience. I felt like I could grasp the basic mechanics after reading the paper, but I wish the documentation could help someone like me be more intelligent with the 'tuning' of the model. I could never get accuracy below 15% average error, which is too large for my use case.
Probably user ignorance, but that's my experience.
edit: found it https://www.reddit.com/r/MachineLearning/comments/pe1lst/r_i...
Turns out it was about time series anomaly detection, but if you can detect, you can forecast if your model is generative
> As an example, let’s look at a time series of the log daily page views for the Wikipedia page for Peyton Manning. We scraped this data using the Wikipediatrend package in R. Peyton Manning provides a nice example because it illustrates some of Prophet’s features, like multiple seasonality, changing growth rates, and the ability to model special days (such as Manning’s playoff and superbowl appearances).
In our vmanomaly product, Prophet is one of the go-to models for anomaly detection in metrics data and it usually requires little tuning to achieve considerable results. The main purpose for the use of Prophet or similar forecasting models is to reformulate the task of anomaly detection:
- given fitted model M, ground truth Y_i for particular data point X_i, we produce forecast Yhat_i and its uncertainty estimate [Yhat_lb, Yhat_ub] - if ground truth Y_i falls beyond the range of [Yhat_lb, Yhat_ub], we consider this point an anomaly - the further Y_i is from the range, the higher the anomaly score would be. In our particular implementation for easier alerting purposes, anomaly_score > 1 means "anomaly"
here's a small visual example: https://docs.victoriametrics.com/vmanomaly.html#examples
“You can imagine my disappointment when, out-of-the-box, Prophet was beaten soundly by a ‘take the last value’ forecast.”
"I'm pretty good at statistics and can predict things using software... I bet I could make money in the stock market"
And then they realize just how hard it is.
in fact, this sort of alternate data is pretty commonplace in firms I've worked at.
VTSAX and chill? :^)
First: Prophet is not actually "one model", it's closer to a non-parametric approach than just a single model type. This adds a lot of flexibility on the class of problems it can handle. With that said, Prophet is "flexible" not "universal". A time series of entirely random integers selected from range(0,10) will be handled quite poorly, but fortunately nobody cares about modeling this case.
Second: the same reason that only a small handful of possible stats/ML models get used on virtually all problems. Most problems which people solve with stats/ML share a number of common features which makes it appropriate to use the same model on them (the model's "assumptions"). Applications which don't have these features get treated as edge-cases and ignored, or you write a paper introducing a new type of model to handle it. Consider any ARIMA-type time series model. These are used all the time for many different problem spaces, and are going to do reasonably well on "most" "common" stochastic processes you encounter in "nature", because its constructed to resemble many types of natural processes. It's possible (trivial, even) to conceive of a stochastic process which ARIMA can't really handle (any non-stationary process will work), but in practice most things that ARIMA utterly fails for are not very interesting to model or we have models that work better for that case.
Out of all possible inputs, there are some that the model works well on and others that it doesn't work well on. The trick is devising an algorithm which works well on the inputs that it will actually encounter in practice.
At the obvious extremes: this library can probably do a great job at predicting linear growth, but there's no way it will ever be better than chance at predicting the output of /dev/random. And in fact, it probably does worse than a constant-zero predictor when applied to a random unbiased input signal.
Except that it's also usually possible to detect such trivially unpredictable signals (obvious way: run the prediction model on all but the last N samples and see how it does at predicting the final N), and fall back to a simpler predictor (like "the next value is always zero" or "the next value is always the same as the previous one") in such cases.
But that algorithm also fails on some class of inputs, like "the signal is perfectly predictable before time T and then becomes random noise". The core insight of the "No Free Lunch" theorem is that when summed across all possible input sequences, no algorithm works any better than another, but the crucial point is that you don't apply signal predictors to all possible inputs.
Another place this pops up is in data compression. Many (arguably all) compressors work by having a prediction or probability distribution over possible next values, plus a compact way of encoding which of those values was picked. Proving that it's impossible to predict all possible input signals correctly is equivalent to proving that it's impossible to compress all possible inputs.
Another way of thinking about this: Imagine that you're the prediction algorithm. You receive the previous N datapoints as input and are asked for a probability distribution over possible next values. In a theoretical sense every possible value is equally likely, so you should output a uniform distribution, but that provides no compression or useful prediction. Your probabilities have to sum to 1, so the only way you can increase the probability assigned to symbol A is to decrease the weight of symbol B by an equal amount. If the next symbol is A then congratulations, you've successfully done your job! But if the next symbol was actually B then you have now done worse (by any reasonable error metric) than the dumb uniform distribution. If your performance is evaluated over all possible inputs, the win and the loss balance out and you've done exactly as well as the uniform prediction would have.
We ran it on such a dataset and found out that directly using https://github.com/karpathy/minGPT consistently gives a better result. So we ended up using the output of Prophet as an input feature to a neural network, but the result was not improved in any significant way.
If anyone is not aware there are many periodic phenomena in astronomy - e.g. variable stars which can have periods from minutes to hundreds of days.
The description of this library sounds like it's very tied to the human world - talking about yearly, weekly and daily seasonality.
[Weirdly though, we do sometimes see variability on 'human' timescales in astronomical data series. If maintenance is carried out weekly on a Monday that can add a signal into the data through missing datapoints.]
If you are looking for an actual timeseries method I would checkout either darts [0] or statsforecast [1]. They are currently the most mature timeseries packages.
[0] https://unit8co.github.io/darts/ [1] https://github.com/Nixtla/statsforecast
If I could build it again, I’d start with automating the evaluation of forecasts. It’s silly to build models if you’re not willing to commit to an evaluation procedure. I’d also probably remove most of the automation of the modeling. People should explicitly make these choices.
Having worked on similar Bayesian time-series forecasting tools at Google, this matches my experience (though I've never used Prophet seriously, so please don't take this as any direct judgement of it as a software package). There is a lot of value in a framework that lets you easily experiment with different model structures (our version of this was the structural time series tools in TensorFlow Probability, see, e.g., https://blog.tensorflow.org/2019/03/structural-time-series-m...). But if you're forecasting something you actually care about, it's usually worth the time to try to understand yourself what structure makes sense for your problem, and do a careful evaluation on held-out data with respect to whatever metric you're really trying to optimize. A fully automated search over model structures is cute, but even when it works, it mostly just ends up rediscovering properties of the data you could or should have already known (e.g., of course traffic to your work-related website will have a day-of-week effect), so the cases where it really adds practical value are harder to find than you might like.Even in the age of deep learning, I do think these relatively classical Bayesian models have a lot of value for many applications. Time-series forecasting tends to be a case where:
- you don't have a ton of iid data points (often, only a single time series),
- you'd like forecasts with principled uncertainty estimates, e.g., credible intervals, giving you a range of scenarios to plan for,
- you often do have a pretty good idea of what features are relevant to the process you're predicting, and
- you want to understand in detail what features the forecast is accounting for (and what it might be missing),
all of which play to the strengths of more classical, structured statistical models, compared to more data-hungry black-box deep learning models. So the basic ideas in Prophet and similar tools do still have a lot of relevance going forward, IMHO.
The quality of the uncertainty estimates is a question though.