Using symbolic regression to predict rare events
turingbotsoftware.com
turingbotsoftware.com
I’m sorry if I’ve grossly misunderstood what you’re doing here, but trying to sell this method in a GUI tool (making it usable by people without a background in statistics) seems almost negligent.
https://web.archive.org/web/*/https://turingbotsoftware.com/...
Why think it’s ok to dumb down data science so much? We don’t use “Be your own family’s surgeon” or “Represent your mom in court” apps! Expertise matters in ML/stats too...
But yeah, with such small sample size it might not generalize as much
let F = transaction being fraud
let D = detected as fraud
P(F) = 492/284807 = .17%
P(D|F) = 80%
P(~F) = 1-P(F) = 99.83%
P(D|~F) = 13%
P(F|D) = (P(D|F) * P(F)) / (P(F) * P(D|F) + P(~F) * P(D|~F))
P(F|D) = ((.8) * (.0017)) / ((.0017) * (.8) + (.9983) * (.13))
P(F|D) = 0.01 = 1% chance of the transaction being fraud given a positive detection
although to be fair, they're not stating to only use this method by itselfOtherwise it's impossible to make the comparisons at the end of the article to other results.
I wonder how this performance would have compared to a simple random forest or MLP model.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.394...
OP misses the point that SR is the problem space, not the algorithm or solution space.
You could afterwards use some feature selection, regularization etc. to retain features that have explanatory power.
It wasn't clear to me why symbolic regression would be especially good for highly skewed datasets. Does anybody have an idea?
https://en.m.wikipedia.org/wiki/Base_rate_fallacy#False_posi...
In practice you'll usually want to tune the tradeoff between precision and recall for the situation. This way you can make a direct connection between model performance and the costs and benefits as they manifest in practice, as well as your tolerance for risk.
Is the cost of allowing a fraudulent transaction through much greater than the cost of blocking a non-fraudulent transaction? Then the error function should reflect that. Minimising F1/accuracy/log-loss is not necessarily going to save the most money.