We added a heterogeneous dataframe auto vectorizer to our oss lib last year for a few reasons. Imagine writing: `graphistry.nodes(cudf.read_parquet("logs/")).featurize(**optional_cfg).umap().plot()`
We like using UMAP, GNNs, etc for understanding heterogeneous data like logs and other event & entity data, so needed a way to easily handle date, string, JSON, etc columns. So automatic feature engineering that we could tweak later is important. Feature engineering is a bottleneck on bigger datasets, like working with 100K+ log lines or webpages, so we later added an optional GPU mode. The rest of our library can already run (opt-in) on GPUs, so that completed our flow of raw data => viz/AI/etc end-to-end on GPUs.
To your point... Most of our users need just numbers, dates, text, etc. We do occasionally hit the need for images... but it was easy to do externally and just append those columns. A one-size-fits-most is not obvious to me for embedding images when I think of our projects here. So this library is interesting to me if they can pick good encodings...