Forecasting with Trees (2021)
amazon.science
amazon.science
When I used to follow, until a few years ago, the winning models were ensembles of ensembles (e.g., RF is an ensemble). The fact that the best single models are ensembles, or evolutions of ensembles, is therefore not surprising.
When dealing with numerical data, squeezing blood from the stones, which is what happens in the latter stages of the prediction competition, is very rarely worth squeezing in the real world. When the model is not mechanistic but only correlative (almost all models are not purely correlative or mechanistic, anyway), getting to the last decimal place of mean absolute error or a similar metric requires building an increasingly complex structure over which we have little control upon a building that has its foundation of sand. All it takes is a little wind, such as a change in the distribution of data over time-which always happens-and unstable structures are bound to collapse.
I have about one million rows of tabular data, with 15 features, to make price predictions.
Is there a definitively better choice between the two?
XGBoost is the og and the most feature-rich.
LightGBM is the fastest and what I use for my case (millions of rows of data with over 100 features).
CatBoost could be good depending on the nature of your data, for example if you have a lot of categorical types.
EDIT: Alos, they all support GPU training, but I haven't been able to make that faster than just using more CPU cores.
All three are lightyears ahead of naive random forest implementations, and are in very active development.
https://www.kaggle.com/competitions/m5-forecasting-accuracy/...
Install python 3.11, and the libraries sklearn, lightgbm, pandas, matplotlib and numpy
Ask a LLM to write a python script that loads the data and fits a model to the data and summarizes/plots some results.
Jupyter Lab with autoreload, and using a python virtual environment are recommended.
And here is a practical intro to it, you can run it right in your browser if you open it in Colab: https://www.tensorflow.org/decision_forests/tutorials/beginn...
You sound like a moron.
Here's a simplified version of the approach (i.e. performing strong feature engineering, then converting the multivariate time series data to a panel/tabular dataset and training a boosting trees model on it), using temporian (a much improved alternative to pandas for working with temporal data) and xgboost: https://temporian.readthedocs.io/en/stable/tutorials/m5_comp...