PREDICTING LOAN APPROVAL USING ML ALGORITHMS
DOI:
https://doi.org/10.64751/r3str357Abstract
Banks and finance companies receive a large number of loan applications every day, and each application must be examined to decide whether the applicant is likely to repay. The decision involves checking income, employment, existing obligations, credit history, the value of any collateral, and the purpose of the loan. When this assessment is performed manually, it takes time, depends on the experience of individual officers, and can produce inconsistent outcomes for applicants with similar profiles. This paper presents a system that predicts loan approval using machine learning algorithms, with the aim of supporting faster and more consistent preliminary decisions. The system uses publicly available loan application datasets containing applicant details such as gender, marital status, number of dependants, education, employment type, applicant and co-applicant income, loan amount, loan term, credit history, and property area, together with the final approval status. The raw data contains missing values, skewed income distributions, and categorical fields that must be converted into numerical form. A preprocessing stage imputes missing values using suitable strategies for each field, applies log transformation to skewed amounts, and encodes categorical variables. Derived features such as total household income, equated monthly instalment, and the ratio of instalment to income are created because they reflect repayment capacity more directly than the raw values. Several classification algorithms are trained and compared, namely Logistic Regression, K-Nearest Neighbours, Naive Bayes, Decision Tree, Random Forest, Support Vector Machine, Gradient Boosting, and XGBoost. The imbalance between approved and rejected applications is handled with class weights and oversampling of the minority class. Hyperparameters are tuned using stratified cross-validation, and the models are evaluated on a held-out test set using accuracy, precision, recall, F1- score, and ROC-AUC. The results show that Random Forest and Gradient Boosting achieve the best overall performance, and that credit history, the instalment-to-income ratio, and total income are the most influential features. The selected model is deployed through a web application built with Flask, where a user enters applicant details and receives a predicted decision along with the probability of approval and the main factors that influenced the result. An administrator view shows summary statistics of processed applications, the distribution of predicted outcomes, and the performance of the model on recent data. The system is designed to assist loan officers by providing a quick and consistent first assessment, reducing processing time and highlighting applications that need closer review. The final decision remains with the officer. Future work includes fairness checks across applicant groups, the use of richer credit bureau information, and integration with digital document verification.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







