HireShield: Deep Learning Based Detection of Fraudulent Online Job Postings
DOI:
https://doi.org/10.64751/f06c2p56Abstract
The rapid proliferation of online job portals has been accompanied by a surge in fraudulent job postings, causing significant financial and psychological harm to job seekers globally. This paper presents HireShield, a deep learning-based binary classification system for real-time detection of fraudulent online job advertisements. HireShield integrates a fine-tuned BERT transformer backbone with a hybrid feature engineering pipeline encompassing textual TF-IDF features, semantic sentence embeddings, and structured metadata flags extracted from the Employment Scam Aegean Dataset (EMSCAD) augmented with 4,680 scraped real-world postings (17,880 records total). A SMOTE-balanced training strategy and progressive BERT layer unfreezing address severe class imbalance (approximately 5% fraudulent postings). HireShield achieves a classification accuracy of 96.8%, precision of 95.4%, recall of 94.9%, and F1-score of 95.1%, outperforming all evaluated baselines including Seq2Seq (70.2% F1), Logistic Regression (75.8%), Random Forest (79.1%), LSTM (83.6%), and fine-tuned BERT without hybrid features (87.3%). Deployed as a Flask REST API with sub-second inference latency, HireShield provides a practical and scalable solution for protecting employment platforms and job seekers from recruitment fraud.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.







