Comparison of Data Mining Classification Algorithms Determining the Default Risk

Publication Date

2019-02-03

Country of Publication

Egypt

No. of Pages

Main Subjects

Mathematics

Abstract EN

Big data and its analysis have become a widespread practice in recent times, applicable to multiple industries.

Data mining is a technique that is based on statistical applications.

This method extracts previously undetermined data items from large quantities of data.

The banking and insurance industries use data mining analysis to detect fraud, offer the appropriate credit or insurance solutions to customers, and better understand customer demands.

This study aims to identify data mining classification algorithms and use them to predict default risks, avoid possible payment difficulties, and reduce potential problems in extending credit.

The data for this study, which contains demographic and socioeconomic characteristics of individuals, were obtained from the Turkish Statistical Institute 2015 survey.

Six classification algorithms—Naive Bayes, Bayesian networks, J48, random forest, multilayer perceptron, and logistic regression—were applied to the dataset using WEKA 3.9 data mining software.

These algorithms were compared considering the root mean error squares, receiver operating characteristic area, accuracy, precision, F-measure, and recall statistical criteria.

The best algorithm—logistic regression—was obtained and applied to the real dataset to determine the attributes causing the default risk by using odds ratios.

The socioeconomic and demographic characteristics of the individuals were examined, and based on the odds ratio values, the results of which individuals and characteristics were more likely to default, were reached.

These results are not only beneficial to the literature but also have a significant influence in the financial industry in terms of the ability to predict customers’ default risk.

American Psychological Association (APA)

Çığşar, Begüm& Ünal, Deniz. 2019. Comparison of Data Mining Classification Algorithms Determining the Default Risk. Scientific Programming،Vol. 2019, no. 2019, pp.1-8.
https://search.emarefa.net/detail/BIM-1210766

Modern Language Association (MLA)

Çığşar, Begüm& Ünal, Deniz. Comparison of Data Mining Classification Algorithms Determining the Default Risk. Scientific Programming No. 2019 (2019), pp.1-8.
https://search.emarefa.net/detail/BIM-1210766

American Medical Association (AMA)

Çığşar, Begüm& Ünal, Deniz. Comparison of Data Mining Classification Algorithms Determining the Default Risk. Scientific Programming. 2019. Vol. 2019, no. 2019, pp.1-8.
https://search.emarefa.net/detail/BIM-1210766

Data Type

Journal Articles

Language

English

Notes

Includes bibliographical references

Record ID

BIM-1210766

SaveSaved Print

Arab Citation & Impact Factor "Arcif"

Largest Arabic Database of Citations Analysis for the Arabic Scholarly Journals Issued in Arab World.

eMarefa Indicators
for Arab Scientific Production

"Kashif" for Checking Similarity or Plagiarism in the Arabic Researches. know more