Logistic Regression by Discretizing Continuous Variables via Gradient Boosting
cdn.rawgit.com
cdn.rawgit.com
Noob here but aren't Bayesian Network DAG just a specialize Neural Network? If so you can use Dirichlet Distribution for Bayesian Network and that's discrete... Unless I'm misunderstanding.
Trying to learn discrete rules is harder because the learning procedure uses gradients to adjust parameters, and the gradients will be zero in a lot more places with discrete "rules".
Gradient Boosted Trees are probably the main thing that comes to mind, but they're not really deep learning.
People have tried to learn hard vs soft attention mechanisms, and while hard attention is faster, it results in worse accuracy and is harder to train.
The inference I draw is that most of the things we want to learn are not described well by discrete rules.
However, people's rationales for why they should bin is often that it makes the model better / more interpretable, without actually testing the more restricted binned model against the more general one. There's certainly something to be said for knowing your audience when choosing a model, though :).