Sift Science (YC S11) raise $18M to stop credit card fraud with machine learning
techcrunch.com
techcrunch.com
By that accounting, a legitimate sale would count as a 'lose', since they lose the TV.
'Lose-lose' would be fine, but even by the standards of hastily written press releases, this is kind of silly.
Jason here, op and CEO of Sift Science. You have a point, but do keep in mind that the TV is going to a bad customer -- one that won't reward Best Buy with repeat business (perhaps just more fraud) and won't spread positive word of mouth (except to let other fraudsters know that Best Buy is a great fraud target). So there is some "lose" in shipping the TV to a bad customer, different from shipping it to a good customer. Does that make sense?
No charge for the first 10k transactions/month is impressive.
For online purchases (card-not-present transactions), the merchant takes on all of the fraud risk. This means that the banks do not have much incentive to protect against fraud. The merchants must go on the information they have, which probably doesn't include the card holder's verified phone number.
Sift Science has been successfully protecting online businesses for over a year, and it turns out that machine learning is (unsurprisingly) a good tool for this sort of classification problem.
What about implementing a 2FA system for larger purchases (online or off), implemented in an app on the consumer's phone like google authenticator or sms? Swipe your card at checkout, if amount is > $XX (or otherwise suspicious according to current models), prompt the buyer for a one-time code from SMS or an app. I use the same system when logging into gmail, my bank account, etc - I'd have no problem (and would even welcome) a similar system when using plastic. It's at least a lot more convenient than having the txn declined and your card disabled until you call their security hotline. This way, thieves would need to steal your card and your phone to cause damage.
Edit: FIs also have to follow rules and regulations to monitor and predict fraud activity.
When the actual owner of the credit card charges back, the $1000 BestBuy received is automatically taken from their account (or subtracted from their next settlement).
FIs do try and predict fraud, but they do a very poor job of it. In the ecommerce space, the merchant is liable for, essentially, all fraud.
I guess I spoke out of context!
I've contacted the reporter to try and clear this up.
Edit #2 - when you say offline/online, I'm assuming you mean in terms of an online store, versus a physical retail store.
I took it as if the transaction was done in realtime to the backend FI system, or if it were verified offline at a terminal. In EMV terminals at a physical merchant location, transactions can be either online or offline ;)
So the actual question here is whether the ML for detecting fraud is better than a flexible rules engine and goods analysts/statisticians. In my personal experience, statistical analysis and anomalies detection effectively handles majority of the fraud. I would be interested to see a more detailed analysis of Sift Science performance with some numbers for false positives/false negatives for example, though (of course) it is probably proprietary information.
Kudos to the Sift Science team!
Many Credit card fraud prevention systems such as the Falcon Fraud Manager use both Rules and ML (Neural Network modelling) to tackle this issue. I am not sure that a purely data centric approach with no rules even makes sense.
but, rules are rather easy for fraudsters to circumvent, and they require merchants to play whack-a-mole. with today's technologies, it's easier than ever to analyze massive amounts of data, and we believe that machine learning can go a really long way in detecting fraud.
does that make sense? happy to discuss further, and we'd be happy to put you in touch with our customers if you'd like to hear more about our results.
Having worked in this space a bit, I know where the dragons be in such an approach. I imagine that you are doing some kind of Pattern Matching: e.g. known Fraudster uses these signals (Browser +OS + Email+Time of shopping + something else... + type of Card) to teach the system what to look for. The real trick is to avoid over training and evolving the patterns to keep pace with the fraudsters by incorporating feedback from the merchants.
Can you point to some blog posts/text that provides a sneak peak into the kinds of technology that you use to cut down on false positives?
Fraudsters generally follow fairly specific patterns that can definitely be picked up in a rich enough dataset, and the consortium-based approaches of Falcon and Sift allow the models to generalize pretty well. Rules are way less expressive than complex-enough machine learning methods (combined with good data).
Rules composed into cascdes or manual decisions trees are not as powerful as ml.
But rules can be converted into features, which you use as primitives in an ml model.
In fact, one pattern for feature engineering is to build a good rules-based system. You then can treat that as features of a degenerate model, with weights of plus or minus infinity. By training, you can induce better weights.
Edit: in a sense, rules and training method are almost orthogonal concerns.
"Rule" has a very specific meaning in this industry, and it's basically a manual decision that exists outside the ML model - usually written by experts. The people who build the models never see the rules (per se), and the rules are integrated after the models are deployed. I wouldn't describe a modeler integrating logic into the modeling process a "rule" because it would confuse terminology.
The whole rule-based thing is sort of a testament to how old-fashioned the industry is IMO. The problem is two-fold: banks don't fully trust ML (or the "experts" want to keep their jobs), and they don't tolerate much change in their computing platforms. That last point is pretty important: banks/issuers don't refresh their fraud models very often, so they can miss new types of fraud, which ends up being a great use-case for rules.
I have been integrating it for a day or so. The documentation is slightly confusing and they've had a few minor bugs in their UI, but support has been really good thus far.