Nice writeup. It seems like a supervised learning approach to fraud detection. I have a question: Where does the is_fraud variable come? Is it done by humans?
Yes, this variable is usually set after the fact. For instance, a given transaction may have led to a chargeback, or may be done by a known fraudster. These models are usually trained on historical data, so we can know with some certainty which transactions are fraud.
It could be that a transaction was fraudulent but has not yet led to a chargeback (maybe the real cardholder hasn't yet seen their statement?), so there's still some uncertainty, but hopefully that approaches a minimum after some time passes.