Show HN: AutoML Python Package for Tabular Data with Automatic Documentation
github.com
github.com
I am knee deep in a personal project exploring machine learning on tabular data and it’s been consuming my off-hours brain for a while. And I pop open HN on a holiday Monday to find…a package for machine learning on tabular data :)
Curious if anyone has any other suggestions of frameworks or packages to explore. It seems that the state of the art in tabular data got a lot of activity in 2019 and 2020 and the industry’s focus moved on to image processing (DALL-E, Stable Diffusion) so I’m wondering whether there’s been much advancement.
What is more, MLJAR AutoML is checking much simpler algorithms for you, like Dummy Models (average response or majority vote), linear models, simple decision trees - because very often you don't need Machine Learning. The xgboost/lightgbm cant do this :)
You can pass the name of the columns that are categoricals when constructing the Booster or Data, and LightGBM will work with them under the hood treating them as a 1 hot encoded.
LightGBM also has a way of automatically treat missing values as either zeros, their own category, or the sample average (I might be mistaken on the last one)
All in all, you still need to do feature engineering and the like, but LightGBM removes a lot of the hassle from Xgboost.
One newbie question maybe you can answer: Can XGboost/LightGBM handle out of band data predictions? The specific regression problem I’m tackling involves price predictions on tabular data (similar the the kaggle housing price problem) and I know classic random forest / decision trees have trouble with time series predictions. Not sure if those models handle that better.
This is not true. The trick is you have to convert your longitudinal data to be cross-sectional via feature engineering of lagged features. There are also related tricks like expanding datetimes into features like day of week, day of month, etc. This can be a lot of work, though there are software tools which help do this.
Some general time-oriented feature engineering plus a vanilla random forest is a great second baseline (after LOCF), and then if needed you can spend 10x the time tuning a GBM to beat that.
If you have a GPU, checkout fastai's tabular learning stuff; easy feature eng and neural nets with embeddings can do a lot with low effort.
As part of improving autoML in PyGraphistry, we have been a doing a lot of auto-feature-engineering and auto-clustering work for one-liners like:
g.nodes(pd.read_csv(...)).featurize().umap().plot()
It incorporates a variety of cool base libraries like dirty_cat and transfomers, adds some of our own, and is heavier on text feature column support than most packages here. Think improving upon TopicBERT or dirty_cat for real text & multicolumn data and interactively visualizing, all in one line. We are currently adding end-to-end GPU support for that pipeline (the rest of our stack already supports that), primarily around cuml GPU data frame support.
It powers a lot of our new visual no-code AI layers and is driven by a bunch of enterprise projects in areas like cyber, fraud, social, misinfo, supply chain, & finance :) More niche, this is all feeding into our automatic graph ai (GNN) layers. Graph AI packages can't truly do graphs well until they can do node tables and edge tables well on their own!
When I managed a deep learning team at Capital One, my last technical project was automatic deep learning architecture search. I think that the fields of data engineering, data science, machine learning, and deep learning are all all ripe for massive automation, reducing the number of jobs in these fields. I think that there will be accelerated use of all of these fields, just much less manual work will be required.
[0] https://sebastianraschka.com/blog/2022/deep-learning-for-tab...
In case, you wanna check it out: https://github.com/microsoft/FLAML