Figure 7 of the mentioned paper evaluates FB-Prophet in a extremely convenient environment of long horizon h in {30,60,90,120,150,180}. It is known that ETS and ARIMA models concatenate errors and degrade in performance with longer forecasting horizons. We have explored and offered solutions to these issues with the N-HiTS model specialized in long-horizon (
https://arxiv.org/abs/2201.12886).
In recent years, FB-Prophet has gained a reputation for the poor quality of its predictions in many practical scenarios (short/medium term horizon) and its slow performance on bigger data sets; our ARIMA/ETS work and NeuralProphet confirmed those suspicions. The mentioned 92 percent improvements of the paper are restricted to h in {1,3,15,60} (https://arxiv.org/pdf/2111.15397.pdf).
This post's results are rather for short-horizon tasks, the same as NeuralProphet experiments. But we are confident that specialized tools like N-HiTS would outperform Prophet in long-horizon settings.