HNHacker News
TopNewBestAskShowJobs

quantperson

4 karma · joined December 6, 2021

submissionscomments
quantperson··on Automated Time Series Processing and Forecasting
It's not hard to beat Prophet by >20%. As the paper itself states, it's designed for rapidly producing good enough forecasts at scale by domain experts with no stats background. If you have a small team consisting of domain experts, engineers and data science/statisticians, you should be able to build something which handily beats Prophet. If you're careful, you should be able to design a pipeline which can do the model tuning and selection in something resembling an "automated" manner (but it requires extreme care, and Prophet can't save you from the dangers herein either).

Prophet has a relatively narrow use case where it makes sense; not coincidentally, it makes sense at a place like Facebook where the sheer variety of problems to be tackled cannot scale to such a foregoing team for each such problem.

Prophet has received a lot of backlash for workflows it never intended, nor claimed, to be able to tackle. That's undeserved in my opinion. Zillow didn't lose money because it used Prophet, for example. But it still really shouldn't be used as a benchmark like this.

quantperson··on Automated Time Series Processing and Forecasting
Great, another open source tool purporting to solve time series analysis in an "automated way" that my manager will link me tomorrow and ask me to review (as an aside, attempting automated statistics of any kind is incredibly dangerous and misguided, but especially so for time series).

Why should I use this over Darts[1] or just Statsmodels[2], if I need more lower level access and diagnostics? Both of these are far more established.

I dislike that Facebook Prophet was chosen as a benchmark; it's not a difficult benchmark to beat for the majority of time series use cases. It signifies to me that this project might targeting cargo cult data science. Prophet is not particularly good at non-daily timeseries and non-seasonal timeseries. The paper itself admits this[3]. Moreover, it's just a generalized additive model that incorporates holidays.

I don't intend to sound demeaning here, really. But I'm trying to understand what the point is. This doesn't look like someone's weekend project, but we already have plenty of established projects which tackle this effectively.

There are three major markets for time series work:

1. You're an analyst with a lot of domain knowledge who needs to analyze daily, seasonal data but you don't have a strong statistical or engineering background. This person should probably just choose Prophet (again, the developers of Prophet explicitly acknowledge that it's designed for scalable good enough models by non-stats people, not for the best model given the data).

2. You're a data scientist with a good statistical background and you need to produce forecasts. You can afford to dig into what the model is doing and select a model based on a series of diagnostics and knowledge about the data itself. This person should probably choose a more complete suite, like Darts. The important thing here is developing good models quickly while being able to do more than just press a button.

3. You're a data scientist (or statistician) which a very strong statistical background who needs to produce the best model they can for answering a specific question. This person is probably going to use R, Stan, Statsmodels or PyMC to come up with something bespoke. They may or may not need to systematize it, but they don't need to produce quantity over quality.

How does this thing improve the state of the art for any of these markets?

--

1 https://github.com/unit8co/darts

2 https://www.statsmodels.org/stable/index.html

3 https://peerj.com/preprints/3190/