New Explainable AI Algorithms
wagtaillabs.com
wagtaillabs.com
You're right that the way we typically train Tree Ensembles creates a massive number of rules, the walk through Random Forest has more than 100,000 leaves per Decision Tree. Once we start grafting it the number of rules starts to vastly outnumber the amount of training data.
I have some follow up articles planned that will cover this in more detail, but the short answer is that I feel that we often jump to overly complex models up front without fully considering whether the accuracy/complexity tradeoff it worth it. Using Amalgamate I showed how I could have the number of rules without significantly increasing validation error (+5%). I believe that if we're careful using model sophisticated techniques (i.e. Boosting and dense/fully connected/tabular neural networks) then we should be able to create reasonably accurate models that are reasonably straight forward to explain.
If you had a mechanism to subscribe for future updates, I'd do it.
Either way, this is cool stuff.
I've been trying to work out where to put updates, so far I've been using GitHub & Twitter (both @wagtaillabs). I'll keep posting to HN as well (I just had two orders of magnitude more traffic than any other day).
I'd be more than happy for any suggestions on places where people could follow (I've thought about an email list, but I'm not sure how many people actually read emails any more).
(While I have you, would you mind adding an email address to your profile that we can contact you at? We do that sometimes when we want to invite a repost.)
Yep, I've added an email address now :)
I'm excited to play with the code.
Yep, code's on GitHub. I'm more than happy to collaborate, there's heaps of things that need to be done.
https://www.aaai.org/Papers/Workshops/1999/WS-99-06/WS99-06-...
But, I apologize ! It's a bit pimped up compared to the one liner above, I think step 7 in section 4.3 is what I was thinking of :) I did laugh when I dug it out, as I have been working on the first bullet in the conclusion this week!
which was a precursor to the model distallation work from Geoff Hinton: https://arxiv.org/abs/1503.02531
http://www.clungu.com/Distilling_a_Random_Forest_with_a_sing...