Machine learning based intelligent framework for data preprocessing

Joint Authors

Sarwar, M. Suhayl
Sufyan, Muhammad
al-Qayyum, Zia
Abd al-Kalim

Source

The International Arab Journal of Information Technology

Issue

Vol. 15, Issue 6 (30 Nov. 2018)6 p.

Publisher

Zarqa University

Publication Date

2018-11-30

Country of Publication

Jordan

No. of Pages

6

Main Subjects

Information Technology and Computer Science

Abstract EN

Data preprocessing having a pivotal role in data mining ensures reduction in cost by catering inconsistent, incomplete and irrelevant data through data cleansing to assist knowledge workers in making effective decisions through knowledge extraction.

Prevalent techniques are not much effective for having more manual effort, increased processing time, less accuracy percentage etc with constrained data volumes.

In this research, a comprehensive, semi-automatic pre-processing framework based on hybrid of two machine learning techniques namely Conditional Random Fields (CRF) and Hidden Markov Model (HMM) is devised for data cleansing.

Proposed framework is envisaged to be effective and flexible enough to manipulate data set of any size.

A bucket of inconsistent dataset (comprising of customer’s address directory) of Pakistan Telecommunication Company (PTCL) is used to conduct different experiments for training and validation of proposed approach.

Small percentage of semi cleansed data (output of preprocessing) is passed to hybrid of HMM and CRF for learning and rest of the data is used for testing the model.

Experiments depict superiority of higher average accuracy of 95.50% for proposed hybrid approach compared to CRF (84.5%) and HMM (88.6%) when applied in separately

American Psychological Association (APA)

Sarwar, M. Suhayl& al-Qayyum, Zia& Sufyan, Muhammad& Abd al-Kalim. 2018. Machine learning based intelligent framework for data preprocessing. The International Arab Journal of Information Technology،Vol. 15, no. 6.
https://search.emarefa.net/detail/BIM-873978

Modern Language Association (MLA)

Sarwar, M. Suhayl…[et al.]. Machine learning based intelligent framework for data preprocessing. The International Arab Journal of Information Technology Vol. 15, no. 6 (Nov. 2018).
https://search.emarefa.net/detail/BIM-873978

American Medical Association (AMA)

Sarwar, M. Suhayl& al-Qayyum, Zia& Sufyan, Muhammad& Abd al-Kalim. Machine learning based intelligent framework for data preprocessing. The International Arab Journal of Information Technology. 2018. Vol. 15, no. 6.
https://search.emarefa.net/detail/BIM-873978

Data Type

Journal Articles

Language

English

Notes

Includes bibliographical references

Record ID

BIM-873978