COMPARATIVE STUDY OF TREE BASED CLASSIFIERS ON IMBALANCED ONLINE TRANSACTION DATA
DOI:
https://doi.org/10.64751/e3kywj51Abstract
Online payments through cards, wallets, and banking applications have grown enormously in volume, and with them the number of fraudulent transactions has also increased. Detecting fraud automatically is difficult because fraudulent transactions form only a tiny fraction of all transactions, often well below one percent. A classifier trained on such data can reach very high accuracy simply by labelling every transaction as genuine, while failing to catch the cases that matter. This paper presents a comparative study of tree based classifiers on imbalanced online transaction data, with the aim of identifying which methods and which imbalance handling techniques give the most reliable detection of fraud. The study uses a publicly available anonymised dataset of card transactions in which the input features are numerical components obtained through principal component analysis, together with the transaction amount and the time elapsed since the first transaction. The target label indicates whether a transaction is fraudulent. The data is examined for duplicates and scale differences, the amount and time fields are scaled, and the data is divided into training and test portions with stratification so that both contain the same small proportion of fraud cases. Five tree based classifiers are compared, namely a single decision tree, random forest, extra trees, gradient boosting, and extreme gradient boosting, with a light gradient boosting model included as a sixth candidate. Each classifier is trained under several imbalance handling strategies, including no treatment, random undersampling of the majority class, synthetic minority oversampling, a combination of oversampling and cleaning, and cost-sensitive learning with class weights. Resampling is applied only within the training folds so that the test data remains untouched. Because accuracy is misleading on data of this kind, the models are evaluated with precision, recall, F1 score, the area under the ROC curve, and the area under the precision-recall curve, supported by confusion matrices and threshold analysis. The results show that boosting ensembles combined with cost-sensitive learning or moderate oversampling give the best balance between catching fraud and avoiding false alarms, while undersampling improves recall at a considerable cost in precision. A dashboard presents these comparisons along with feature importance and a simple interface for scoring new transactions. The study is intended to give practical guidance to developers of fraud detection components in payment systems, and to show clearly how the choice of evaluation metric changes the conclusions that are drawn. Future work includes testing the models on streaming data, examining concept drift as fraud patterns change over time, and adding explanation techniques that help analysts review flagged transactions.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







