wCM based hybrid pre-processing algorithm for class imbalanced dataset

Singh, Deepika; Saha, Anju; Gosain, Anjana

doi:10.3233/JIFS-210624

wCM based hybrid pre-processing algorithm for class imbalanced dataset

Article type: Research Article

Authors: Singh, Deepika^{; *} | Saha, Anju | Gosain, Anjana

Affiliations: USICT, Guru Gobind Singh Indraprasth University, Sector-16, C, Dwarka, New Delhi, India

Correspondence: [*] Corresponding author. Deepika Singh, USICT, Guru Gobind Singh Indraprasth University, Sector-16 C, Dwarka, New Delhi, India. E-mail: [email protected].

Abstract: Imbalanced dataset classification is challenging because of the severely skewed class distribution. The traditional machine learning algorithms show degraded performance for these skewed datasets. However, there are additional characteristics of a classification dataset that are not only challenging for the traditional machine learning algorithms but also increase the difficulty when constructing a model for imbalanced datasets. Data complexity metrics identify these intrinsic characteristics, which cause substantial deterioration of the learning algorithms’ performance. Though many research efforts have been made to deal with class noise, none of them focused on imbalanced datasets coupled with other intrinsic factors. This paper presents a novel hybrid pre-processing algorithm focusing on treating the class-label noise in the imbalanced dataset, which suffers from other intrinsic factors such as class overlapping, non-linear class boundaries, small disjuncts, and borderline examples. This algorithm uses the wCM complexity metric (proposed for imbalanced dataset) to identify noisy, borderline, and other difficult instances of the dataset and then intelligently handles these instances. Experiments on synthetic datasets and real-world datasets with different levels of imbalance, noise, small disjuncts, class overlapping, and borderline examples are conducted to check the effectiveness of the proposed algorithm. The experimental results show that the proposed algorithm offers an interesting alternative to popular state-of-the-art pre-processing algorithms for effectively handling imbalanced datasets along with noise and other difficulties.

Keywords: Classification, class imbalance, data complexity, overlapping, bayes error, pre-processing, learning algorithms

DOI: 10.3233/JIFS-210624

Journal: Journal of Intelligent & Fuzzy Systems, vol. 41, no. 2, pp. 3339-3354, 2021

Published: 15 September 2021

Price: EUR 27.50

North America

IOS Press, Inc.
6751 Tepper Drive
Clifton, VA 20124
USA

Tel: +1 703 830 6300
Fax: +1 703 830 2300
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

Europe

IOS Press
Nieuwe Hemweg 6B
1013 BG Amsterdam
The Netherlands

Tel: +31 20 688 3355
Fax: +31 20 687 0091
[email protected]

For editorial issues, permissions, book requests, submissions and proceedings, contact the Amsterdam office [email protected]

Asia

Inspirees International (China Office)
Ciyunsi Beili 207(CapitaLand), Bld 1, 7-901
100025, Beijing
China

Free service line: 400 661 8717
Fax: +86 10 8446 7947
[email protected]

For editorial issues, like the status of your submitted paper or proposals, write to [email protected]

如果您在出版方面需要帮助或有任何建, 件至: [email protected]

Share this:

North America

Europe

Asia