Tehnički vjesnik, Vol. 30 No. 2, 2023.
Izvorni znanstveni članak
https://doi.org/10.17559/TV-20221119040501
MRMR-EHO-Based Feature Selection Algorithm for Regression Modelling
Sathishkumar V. E.
; Hanyang University, Department of Industrial Engineering, 222 Wangsimini-ro, Seondong-gu, Seoul, Republic of Korea, 04763
Yongyun Cho
; Sunchon National University, Department of Information and Communication Engineering, Suncheon, Republic of Korea
Sažetak
In the classical regression theory, a single function model is fit to a data set. In a complex and noisy domain, this process is too complex and/or not reliable. Piecewise regression models provide solutions to overcome these difficulties. The regression performance can be improved by proper feature selection. This paper proposes a feature selection technique for improving regression problems using the hybridization of filter and wrapper feature selection methods. It uses a hybrid framework of Elephant Herding Optimization (EHO) and minimum Redundancy and Maximum Relevance (mRMR). The mRMR-EHO is implemented to maximize the performance of individual regression algorithms and the results are provided in this research. In this paper, the effectiveness of CUBIST and mRMR-EHO feature selection using six fine grained data from small-sized data to big data is empirically demonstrated such as: a) Strawberry Plants Nutrient water supply, b) Steel Industry Energy Consumption, c) Seoul Bike Sharing Demand, d) Seoul Bike Trip duration, e) Appliances energy consumption dataset, f) Capital Bike share program data the results show a marginal increase in performance even to a very large scale. All 6 datasets were pre-processed well for building the models. The empirical results are based on the following algorithms: a) Generalized Linear Regression, b) K nearest neighbour, c) Random Forest, d) Support Vector Machine, e) Gradient Boosting Machine, f) CUBIST. Their performances are compared, and the best-performing model is selected. Ultimately, this paper puts forth that the mRMR-EHO-based feature selection with the rule-based CUBIST model for regression can be used as an effective tool for predictive data modelling in various domains.
Ključne riječi
data mining; elephant herding optimization; feature selection; machine learning; MRMR
Hrčak ID:
294390
URI
Datum izdavanja:
26.2.2023.
Posjeta: 1.210 *