Background
Fraud detection modeling is often considered a classic binary classification problem. In practice, however, even a classifier with 99% recall or 99% PR-AUC can still be useless. An extremely imbalanced dataset can cause a classifier to produce a huge number of false positives, which will reduce profits from transaction fees.
In this article, I will discuss two major topics that are rarely discussed:
- How to define a proper goal and a measurable metric
- Data preparation
In the next article, I will cover training practices:
- Techniques for handling imbalanced and huge datasets
Goal Setting
Pitfalls
Goal setting is the most critical step in a data analytics project. Spending time discussing it with your stakeholders to figure out the real business goal is one of the best investments you can make. It’s much more important than putting time into modeling. I’ll use my experience building a fraud detection model as an example.
Before discussing what the goal should be, let me first cover two common pitfalls:
-
Treating a fraud detection model as the goal. This is probably the most common mistake: defining the development of a fraud detection model as the goal. However, building a model is just an approach to achieving a specific goal, not the goal itself.
-
Setting ROC/PR-AUC as the target metric. These are useful metrics for evaluating the model. Unfortunately, good ROC/PR-AUC can only indicate how well the model performs as a classifier; it doesn’t tell you how well the model is aligned with the business goal.
Building a fraud detection model and measuring its ROC/PR-AUC are both valid parts of an analytics project. Just don’t mistake them for the goal.
Define the Business Goal
So, what should the goal be? The goal should be the profitability of fraud prevention. Although this sounds straightforward, it is often ignored. It’s easy to think, “Fraud eats into profits, so we have to prevent it.” Nevertheless, fraud prevention is a double-edged sword. There are two consequences:
- Correctly identifying fraudulent transactions (True Positive): it boosts profits by preventing losses from fraudulent transactions.
- Incorrectly flagging legitimate transactions as fraud (False Positive): it reduces profits because the service cannot earn from those transactions.
In other words, fraud prevention can result in greater losses than doing nothing. The key is balancing true positives against false positives.
Quantify the Goal
A payment service provider earns profits from transaction fees and loses profits because of fraudulent transactions. Here’s how I define the goal for fraud detection modeling.
I compare profit when using the model with profit when every transaction is approved. Only blocked transactions make a difference:
- True positive: blocking a fraudulent transaction and avoiding the associated loss.
- False positive: blocking a legitimate transaction and forgoing its transaction fee.
- True and false negatives: the transaction is approved either way, so there is no difference.
$$ \Delta\Pi = \sum_{i=1}^{N} \hat{y}_i \left[ y_i , \ell , a_i - (1 - y_i) , r , a_i \right] $$
or, more simply:
$$ \Delta\Pi = \ell \sum_{i \in \mathrm{TP}} a_i - r \sum_{i \in \mathrm{FP}} a_i $$
| Symbol | Meaning |
|---|---|
| $\Delta\Pi$ | Extra profit from the model |
| $N$ | Number of transactions |
| $a_i$ | Amount of transaction $i$ |
| $y_i$ | $1$ if transaction $i$ is fraud, else $0$ |
| $\hat{y}_i$ | $1$ if the model blocks transaction $i$, else $0$ |
| $\ell$ | Fraud loss ratio: the share of a fraud amount we lose |
| $r$ | Net take rate: the share of a legitimate transaction we keep as profit |
| $\mathrm{TP}$, $\mathrm{FP}$ | True positives and false positives |
Keep in mind that $\ell$ and $r$ vary across payment services, and the net take rate $r$ is not the transaction fee or merchant discount rate (MDR). For example, if the transaction fee is 3%, credit card issuers can usually earn 2% as $r$. However, PayFacs or wallets usually earn below 1%. Therefore, you need to understand how the transaction fee is split and what portion your payment service can earn in order to get an accurate estimate of $r$.
Data Preparation
After defining a measurable goal, there are two other important topics that are rarely discussed:
- What data/transactions are valid
- How to define fraudulent transactions
Valid Data
This is another easy step to overlook. In reality, not all transactions are suitable for training. Almost all modern payment services already have a rule-based fraud detection system in place. Some transactions have already been blocked. We never know whether those blocked transactions are fraudulent or falsely blocked. Therefore, excluding blocked transactions is an essential step.
Label Data
Fraud detection has a natural source of labels: chargeback transactions. However, people usually get caught up in deciding what counts as “real fraud.” Although a chargeback may occur because of the user, it doesn’t matter. Remember that the goal of a fraud detection model is to boost profits, not to define fraud. My suggestion is to treat chargebacks directly as fraud labels. You don’t need to manually select from them. Usually, the model is robust to those outliers.
Summary
This article discusses goal setting and data preparation. Both are usually ignored by data scientists. I agree that neither is particularly exciting, and learning about the project background, business flow, and related systems takes time. However, they form the foundation of the whole project and are much more important than modeling itself. Moreover, how you define the goal also affects how you approach modeling. In the next article, I will discuss modeling techniques for imbalanced data and cost-sensitive approaches.