A cross-border community for researchers with openness, equality and inclusion
An Intelligent Machine Learning-Based Framework for Scam Call Detection Using Textual and Behavioural Features
ID:85 View protection:Participant Only Updated time:2026-07-25 21:37:21 Views:14 Online

Start Time:2026-07-30 15:40

Duration:15min

Session:[S4] Computer Vision and Pattern Recognition [S4-2] Computer Vision and Pattern Recognition

Abstract
Telecom fraud has widespread financial costs, with total damages surpassing USD 38.95 billion in 2023. Traditional blacklist systems are struggling to keep pace with these adaptive fraud techniques. For instance, techniques such as Caller-ID spoofing and dynamic number rotation pose a major challenge. We propose an intelligent multimodal scam call detection framework, which unifies TF-IDF textual features with synthetically constructed behavioural call metadata such as call duration, call frequency, time-of-day, and attempts to call repeatedly. Our framework brings robustness to binary and multi-class classification of call fraud. The performance of three supervised classifiers, namely Logistic Regression (LR), Linear Support Vector Machine (SVM), and Random Forest (RF), is evaluated on a labeled dataset of 5,926 calls, which is split randomly and stratified in a ratio of 80% for the training set and 20% for the test set. The support of the SMOTE technique and the use of a weighting scheme to counteract class imbalance were applied. Linear SVM performed best and provided the following results for binary classification: accuracy = 99%, precision = 0.96, recall = 0.91, and F1-score = 0.94. In the four-class classification scheme, which includes Normal, Bank Fraud, Lottery Scam, and Emergency Scam, the linear SVM achieved a 98% weighted average performance. Among the behavioral features, call duration (0.081) and call frequency (0.060) ranked highest and found to be the most important and discriminative call features, while the textual tokens ‘free,’ ‘claim,’ and ‘reply’ were the most important textual featuresof the framework. The proposed study demonstrated an improvement over the performance of deep learning baselines, and while preserving interpretability for real-time mobile applications. LSTM, CNN, and BERT techniques were the baselines used.
Keywords
Speaker
ABDUL KHADAR SHAIK
B. S. Abdur Rahman Crescent Institute Of Science And Technology

Post comments
Verification Code Change Another
All comments
Important Dates
  • Conference date

    07-30

    2026

    -

    08-01

    2026

  • 07-28 2026

    Draft paper submission deadline

  • 07-28 2026

    Registration deadline

Sponsored By

The United Societies of Science

Organized By

Kongunadu College of Engineering and Technology

Contact info
×

USS WeChat Official Account

USSsociety

Please scan the QR code to follow
the wechat official account.