I think the problem arises from the datasets used to evaluate the performance of the models. In the case of Prophet's paper, only one time series is used (The number of events created on Facebook). We can conclude from the results comparing AutoARIMA vs. Prophet (https://github.com/Nixtla/statsforecast/tree/main/experiment..., using the same datasets as in the ETS vs. NeuralProphet experiment) that ETS is also better than Prophet. Regarding NeuralProphet vs. Prophet, the results are not conclusive for these datasets.