Arima Model – Guide to Time Series Forecasting in Python
machinelearningplus.com
machinelearningplus.com
and she has an upcoming oreilly book on TS analysis. Preorder now! (no AF, just a fan of her) https://www.amazon.com/Practical-Time-Analysis-Prediction-St...
One point of critique though which is not immediately mentioned in the article: Depending on your data, and especially if your granularity is measured in anything under minutes, you might find the need to continuously calibrate your parameters as to ensure performance over time. I found myself running optimisations in parallel when deploying an ARIMA-inspired model for my use case, for instance. If continuously computing optimisations is not an option, it might be worthwhile implementing outlier detection or simple thresholds to prompt optimisation calls when needed. In either case, ARIMA is definitely a fun rabbit hole to explore if you have not previously worked with time-series forecasting!
For time series in the past I’ve used Bayesian structural time series with good results. Good article comparing ARIMA to bsts here:
https://multithreaded.stitchfix.com/blog/2016/04/21/forget-a...
ARIMA can do seasonality too if you differentiate for it. I believe both require you to either decompose your time series and figure out if there is any seasonality in it. ARIMA requires that the time series be stationary so if it is not you need to transform it. Exponential smoothing including Holt Winters don't care and iirc you use exponential smoothing technique most of the time for nonstationary data.
GLM have rigorous assumptions for data that makes it usually harder for time series. Both ARIMA and Holt Winter are for forecasting where was GLM is for inferences. ARIMA is auto regressive where it regress on it's pass self. So if you look at the equation it's only Y's unless you use ARIMAX. GLM is a general framework for regression and it regress on other predictors (independent variables). It's tricky to do GLM on time series data because when you want to predict a value in the future say 2021 then often time you need to know your predictor values in the future to predict your response.
I know ARIMA inference is base on time lag. Hence regressing it past self. The relationship that your inferencing on is what response you care about and it's past self. So you can detect patterns in time lag. As for GLM you inference between the linear association between the independent predictors and the response.
For example, here I create and train a model:
model = ARIMA(df.value, order=(1,1,1))
fitted = model.fit(disp=0)
And then I immediately do forecast: fc, se, conf = fitted.forecast(...)
Yet, it is not what I need. Typically, I store the model and then apply it to many new datasets which are received later.sklearn explicitly supports this pattern (train data set for training and then another data set for prediction).
Is it possible in ARIMA and other forecasting algorithms in statsmodels?
Overall...If you're just learning time-series analysis, ARIMA is an important lesson to learn, but for actually doing the analysis/forecast, I'd say use Prophet. Just don't fall into the trap of using a package someone else created without understanding what's going on underneath
For textbook problems like forecasting electricity demand, sARIMA works just fine, but you'll struggle forecasting anything more complex. This is not to say that it's impossible, but it will require far more expertise.
In effect, Icinga (Nagios) represents a local maxima. It's very easy to write checks without mathematical or statistical sophistication -- write a check in the language of your choice and return a status enum containing one of three values: OK, WARN, CRIT. A result there's a massive collection of checks written by others. Sure the language to configure it sucks, but mostly because it models so many things -- systems to be monitored, how to group systems, the checks to monitor systems with, who to alert, and on what schedule.
Instead of doing one thing well, Nagios does everything, kinda poorly. You can't easily schedule overrides for an on-call rotation. Event correlation is entirely manual. You have to restart the service to add or remove monitoring -- it's not prepared for autoscaling or clustering. It can't even scale itself!
There is a better way, using a layered approach. Break the task up into multiple steps: Sense, Store, Analzye, Alert, Route. Nagios effectively discards the Store step, and divy up the remaining work between checks and nagios master. Sense+Analyze+Alert are done by the client, and the Nagios master handles alert Routing.
(taking it home to the topic) Because Nagios has no Store step, analyze is limited in power. You cannot calculate standard deviations, or even rates of change. ARIMA is impossible, as is ETS. You typically define static thresholds. If there's any seasonality, you have to find a static threshold that catches real problems while minimizing false alarms.
The downside to the layered model is that there are many different solutions, and every decision point leads to a more fragmented community. How many people out there run statsite+graphite+grafana+opsgenie? A good layered approach needs to fully isolate layers, so that any integration feature or problem is limited to pairings. Usually easy for features, but for bugs, it's almost definitionally impossible to prevent a problem in an earlier layer from causing symptoms or or two layers later.
tl;dr: I've been mostly learning about ETS (holtwinters) on the job, but would love to learn more about predicting customer traffic.