Background
In the last article, I explained how to define a measurable goal and prepare the data. Building a fraud detection model brings its own challenges. Fraudulent transactions make up a tiny minority. Although fraud rates vary across payment services, they are often well below 1%. For example, EBA’s 2025 Report on Payment Fraud, pp. 12 and 15 reports that, in 2024, approximately 0.015% of card transactions and 0.011% of e-money transactions were fraudulent. Class imbalance is the first challenge. Transaction volume adds another: payment services commonly handle over one billion transactions per year.
For reference, Americans made 60.2 billion transactions; Visa handled 323 billion transactions in 2025 according to CapitalOne Shopping Research.
In this article, I share my experience building a fraud detection model in practice. Through a case study, I explain how I handled an extremely imbalanced dataset with a huge number of transactions, discuss common approaches, and describe the resulting pipeline.
Case Study
My dataset came from a rule-based fraud detection system and contained over one billion valid transactions for modeling. The system already had sophisticated rules that filtered out most fraudulent transactions, leaving around 10 thousand that passed the rules. This was an impressive result for a rule-based system, but it left me with a fraud rate of around 0.001% and an especially difficult class imbalance.
The dataset presented two main challenges:
- Extremely Imbalanced Dataset
- Huge Dataset
Choosing the model
I would usually start with data preparation, but two constraints led me to choose the model first: a strict SLA on inference time and an extremely imbalanced dataset. Tree-based models handle imbalanced datasets well, but most take longer to run inference. LightGBM is the only exception. Although it is slower than logistic regression or Naive Bayes, it was fast enough to meet the SLA.
Define a Machine Learning Goal
The machine learning goal defines what the model optimizes and guides hyperparameter tuning. It does not always match the business goal. Learning-to-Rank (LTR), for example, usually optimizes nDCG, while the business goal is more likely to focus on increasing the number of interactions or the average time spent on the platform. Those goals cannot be measured using existing data.
In my case, extra profit could serve directly as the machine learning goal because we could estimate it from historical data. We knew which transactions were fraudulent and the amount of each transaction.
Downsampling and Upsampling
Downsampling legitimate transactions is necessary because they far outnumber fraudulent transactions, and training directly on one billion transactions requires huge computing resources. With hundreds of candidate features, I started with feature selection. I downsampled legitimate transactions to match the number of fraudulent transactions and used a tree model to select features. This left 20 features, allowing me to include more transactions in training. You can learn more about feature selection from sklearn.
After feature selection, I kept all fraudulent transactions and retained 1/1000 of the legitimate transactions, leaving 1M legitimate transactions for training. The available computing resources mainly determined how many I could include. I also tried upsampling fraudulent transactions with SMOTE, but all metrics, including extra-profit and ROC-AUC, were worse than with downsampling legitimate transactions alone. I suspect there are a few reasons:
- Upsampling also amplifies noise.
- SMOTE may benefit weaker classifiers more than a model such as LightGBM in this setting.
Cost-Sensitive Learning or Recalibration
Cost-sensitive learning is a well-known approach to fraud detection modeling. It offers a sophisticated way to account for costs: apply the weight $\frac{1}{\gamma}$ to fraudulent transactions. In my case, however, legitimate transactions had been downsampled to $\frac{1}{1000}$, so they would receive 10x the weight of fraudulent transactions instead: $\frac{w_0}{w_1} = \lambda \times \gamma = 1000 \times 0.01 = 10$, assuming a fraud loss ratio of $\ell=1$. Here, $w_0$ is the weight for legitimate transactions, $w_1$ is the weight for fraudulent transactions, and $\lambda=1000$ is the inverse sampling rate for legitimate transactions. This created a practical problem: the model failed to learn to identify fraudulent transactions and instead predicted every transaction as legitimate.
Instead, I trained a model without additional weights, recalibrated its probabilities to account for downsampling, and applied a threshold based on the net take rate $\gamma$. LightGBM handled the imbalance well even though legitimate transactions were still 100x as numerous as fraudulent transactions.
Pipeline
My training pipeline had four stages:
- Downsampling legitimate transactions
- Training LightGBM without additional weights
- Recalibrating predicted probabilities
- Applying a threshold based on the net take rate $\gamma$
The threshold example in the diagram uses a net take rate of 1% and a fraud loss ratio of 1.
I kept all fraudulent transactions ($y=1$) and randomly sampled legitimate transactions ($y=0$) at a retention rate of $s=\frac{1}{1000}$, leaving 1 million legitimate transactions. I trained LightGBM on this data without additional class weights.
LightGBM’s output $q(x)$ estimates the fraud probability in the sampled data. The recalibration step converts it to the original population’s fraud probability $p(x)$:
$$ p(x)=\frac{s q(x)}{1-q(x)+s q(x)} =\frac{q(x)}{1000\left(1-q(x)\right)+q(x)} $$
This correction accounts for the change in class proportions. It assumes random sampling within each class and accurate probability estimates for the sampled population.
The classifier then applies the profit-based threshold, where $\gamma$ is the net take rate and $\ell$ is the fraud loss ratio:
$$ \tau=\frac{\gamma}{\ell+\gamma}, \qquad \hat{y}(x)=\mathbf{1}\left[p(x)>\tau\right] $$
Here, $\hat{y}(x)=1$ means block, and $\hat{y}(x)=0$ means approve. When $\ell=1$, the threshold is $\tau=\frac{\gamma}{1+\gamma}\approx\gamma$ for a small net take rate. Sampling affects only the preparation of training data. At prediction time, each transaction passes through the classifier and recalibration steps.
For the illustrative net take rate $\gamma=0.01$ and fraud loss ratio $\ell=1$, the threshold becomes:
$$ \tau=\frac{0.01}{1+0.01}\approx 0.009901 $$
That is approximately 0.9901%. Two example scores show how recalibration changes the final decision:
| Sampled fraud probability $q(x)$ | Recalibrated fraud probability $p(x)$ | Prediction at the example threshold |
|---|---|---|
| 50% | $\frac{0.5}{1000(1-0.5)+0.5}\approx 0.000999$, or 0.0999% | Legitimate ($0$): approve |
| 95% | $\frac{0.95}{1000(1-0.95)+0.95}\approx 0.018646$, or 1.8646% | Fraud ($1$): block |
Hyperparameter
I used $\Delta\Pi = \ell \sum_{i \in \mathrm{TP}} a_i - r \sum_{i \in \mathrm{FP}} a_i$ as the objective for hyperparameter tuning, as described in Quantify the Goal in Part 1. Combining Optuna with cross-validation worked well in my case and was the final step toward achieving positive extra profit $\Delta\Pi$ on the test dataset. This means optimizing the entire pipeline, rather than LightGBM alone.
Summary
If you are adding a fraud detection model to an existing rule-based system, you are likely to encounter the same challenges. Strict rules already block most fraudulent transactions, and the huge dataset requires downsampling. In my case, even with a huge, imbalanced dataset, the pipeline delivered positive extra profits on the test set through downsampling, recalibration, and proper threshold setting. I hope my experience inspires your own approach.