Somewhat off-topic, but has anyone tried training a random forest at LLM scale? Like an RF with millions of trees and billions of branches? My intuition says it could be much more efficient on CPUs, provided it works at all.
You can cluster data using unsupervised random forests and then use these cluster indices as features.