Using Graph Databases to Stop E-commerce Fraud in Real Time
neo4j.com
neo4j.com
A proper fraud detection system has lots of pieces. The one the customer was having trouble with was Cross Reference in real time. Imagine 10k requests per second coming in, you can use cross reference to flag transactions for further analysis or as a partial weight to more fraud detection algorithms.
>User ID
1 account per card per IP, basic stuff.
>IP address
Vip72 is a pretty basic example of a commercial solution, more advanced people either source fresh IPs from bots or scan for SSH tunnels.
>Geo location
While not as trivial to defeat, most "serious" fraudsters either operate their own mule networks or pay for access to one. There's no lack of stupid people.
>A tracking cookie
lol?
>Credit card number
There's no lack of credit cards, even if you're paying $30 for a no-AVS card with full VBV info that still costs less than a TV.
You can solve these problems by adding to the list of identifiers. Off the top of my head: browser fingerprinting, request latency, IP subnets, IP ASNs, hsts super cookies...
Stopping fraud is always going to be a cat and mouse game. The goal is not to eliminate fraud completely, but to make it so difficult for fraudsters to use your service successfully, that they would rather move to the next target. If a fraudster is specifically targeting your service, they will get around any blockades you put in front of them. But you can still make it as hard as possible to circumvent those blockades.
If you deter a high enough percentage of fraudsters, you mitigate your risk of chargebacks. That risk will never be zero. But you should certainly do all you can to minimize it.
"How graph databases stop fraud e-commerce frauds in real-time" - note the use of 'stop' and 'real-time'.
"Fortunately, graph database technology is able to detect the patterns that arise around these e-commerce fraud scenarios and put an end to them in real time"
Again, look at the terms used: 'put an end' and 'real time'.
I am not trying to nit-pick, but what has been discussed is far from being able to put an end to fraudsters. Maybe this could weed out some of the newbie frausters, but that is it.
Disclaimer: I work in the payments industry.
However, with relevancy in a graph we can detect this type of fraud significantly better as we can begin to find relationships that stick out and are uncommonly common to each other.
Services like MaxMind go a very long way in identifying proxies or Tor nodes as well.
Nothing is going to be 100% perfect but the more you can do to make it a struggle for the perpetrator the more they have to decide if it's worth the effort.
The flip side of that is that actively engaged attackers become easier to identify in terms of patterns of behavior and leaving trails to alert authorities.
Most of the more "serious" types use setups similar to FraudFox that defeat systems like these (Although this too causes detectable patterns, new customers with 0 tracking cookies putting in orders less than 30 minutes after initially hitting the site are probably fraud). Don't forget that the Pareto principle applies here, to be able to perform large scale fraud you need lots of cards which requires investment which requires that you're successful.
>Services like MaxMind go a very long way in identifying proxies or Tor nodes as well.
Nobody uses Tor, at best that'll prevent some kids using public proxies (which are worth stopping mind you, but that generally only accounts for a small fraction of the fraud). MaxMind is a joke and struggles against stuff like vip72, and has absolutely no chance against anything more sophisticated than that.
To be honest you could probably replace MaxMind with DNSRBLs and do just fine
>The flip side of that is that actively engaged attackers become easier to identify in terms of patterns of behavior and leaving trails to alert authorities.
Authorities do not give a shit, they might go after card shops but they definitely aren't going to do anything about some guy doing orders to Russia.
(The Tor thing is easy because Tor publishes a list of exit nodes, but some people manage to screw even this up by blocking even the non-exit relays as well!)
The Case Against Specialized Graph Analytics Engines
This paper is a bit unrelated to Neo4j. Neo4j is an ACID-compliant native graph database. It is not a "specialized graph analytic engine."
Neo4j stores the graph data on disk (and caches in memory) as nodes and relationships. After index lookups to find the start points in the graph, all traversals of relationships are done in constant time -- allowing it to scale with linear performance characteristics, regardless of the size of the graph.
The referenced paper cites two main reasons that RDBMS would be better for the analytics use cases: (1) ability to express graph queries in SQL and (2) performance of executing those queries.
(1) Cypher is, like SQL, a declarative language. However, it represents graph constructs in a much more natural way -- "ASCII art for graphs." There's significant praise from developers on the web of the benefits of Cypher for traversing graphs, which is why we decided to open up the language: http://www.opencypher.org/
(2) As Neo4j isn't really intended as an analytics engine, its performance characteristics are not included in this paper. However, (expensive) indexes do not need to be created and maintained for every relationship in Neo4j. Similarly, these indexes do not need to be accessed for traversal (also expensive).
This is intended to stop people from using other peoples credit cards, not their own cards they just applied for as someone else.