Show HN: Unplugg: An automated Forecasting API for timeseries data
unplu.gg
unplu.gg
I do believe we have some similar feature in the pipeline for development, I'll make sure to push it forward. Thanks for the feedback.
What kind of details would you say can be inspected to see if the model is reliable? AR or MA orders, inferred seasonalities? They can give me some notion of what kinds of assumptions were created about my data, but do not guarantee that it will perform :/
For instances where that kind of insight is needed, I don't think our way is the way to go, but rather the use of some forecasting package (R's Forecast or FB's Prophet) and a more exploratory work. But we're looking more at instances where what matters are the forecasted values and not so much the information underneath - automated anomaly detection systems, consumer-facing apps, and along those lines.
I do think the confidence intervals/prediction intervals should be accessible and should probably be adjustable (e.g. 99%, 95%, 80%).
Definitely, this is arguably the most important feature.
I can share that our platform is built on top of ARIMA models, but with a lot of pre-processing work done previously to try and figure out automatically the best parameters to use, as well as a lot of previous hand-tweaking done by ourselves in-house using different datasets (we started out tuning it for forecasting energy consumption, but figured that the resulting models were performing well enough to warrant testing in other domains).
Right now we're opening it up for testing to get more feedback on its performance, so feel free to shoot any more questions or feedback.
dear lord, why? this reminds me of the old "xml binary format" joke:
<byte> <bit>0</bit> <bit>0</bit> <bit>1</bit> <bit>0</bit> <bit>0</bit> <bit>1</bit> <bit>0</bit> <bit>0</bit> </byte>
Even if they were wedded to JSON for some reason, they could have just used a list of observations, like:
[1458000000,63.422235],
That would have cut their data costs in half.
Or just use one of the many existing formats for transmitting time series data. It's not a new topic. https://github.com/mobileink/data.frame/wiki/What-is-a-Data-...
I'm not saying it's ideal, I just think the snark is unwarranted considering how common it is. I just checked InfluxDB and they follow a similar model (even more verbose). https://docs.influxdata.com/influxdb/v1.2/guides/querying_da...
Checked a few more and I believe they're the same - Microsoft IoT, Predix (GE), etc.
"columns": [
"time",
"value"
],
and then the observations as a list of lists: "values": [
[
"2015-01-29T21:55:43.702900257Z",
2
],
[
"2015-01-29T21:55:43.702900257Z",
0.55
],
exactly as i suggested in the "even if they were wedded to JSON for some reason" section of my original explanation.EDIT: But XML payloads are actually a really useful idea, it's going ASAP to the Trello board. :p
but in all seriousness, why the timestamp at all? your examples are all spaced at 3600ms. asking for it implies certain behavior. can you handle heterogeneous interval data? missing data?
However, due to how we model the forecast, it isn't realistic to expect ultra-long term predictions, as eventually the forecast will revert to the mean of the series.
In a more practical note, we have seen good results with forecast windows in between 1/4 and 1/8 the size of the historic data given. So, in your case you could expect between 1-3 months of forecast.
As fate has it, we have no connection to FB's Prophet - we at Whitesmith have been working on unplugg for some time now and decided a few weeks ago that this week we'd share it on some communities to have more people testing it and more feedback. It seems that the folks over at Facebook decided something similar. You know what they say, great minds :p
Joking aside, as intimidating as it might have been to see FB releasing a related tool, we feel that we still fill a different segment. From what I've been reading today, Prophet is a tool tailored for timeseries forecasting with human interaction and input in mind - it can work like a black forecasting box but it seems that it is the most useful when paired with an analyst that can keep looking at the output and tweak the model accordingly. It is _really friendly_ as far as forecasting packages go and trust me, we looked at a fair amount of them. That and the use of ProbProgramming to infer their params is just awesome (I'm a fervent Bayesian at heart).
Unplugg on the other hand, fills the need for a "generic" forecasting tool for uses where you don't want/need much specific tailoring and want a really Plug&Play solution - it's an API that you can call from pretty much everywhere, with no dependencies or specific environments needed (so no need to deploy your own R/Python/Matlab - yikes - environment where your models live and run). One possible use case would be an energy monitoring portal that lives completely client-side and requests forecasts to our API on-the-fly directly from the client.
We are still actively developing and testing different forecasting models - the one running is just the one we feel most confident about - and will be looking at Prophet as a possible alternative (although I haven't seen their license carefully, so can't be sure).