A Classification-Based Machine Learning Approach to the Prediction of Cyanobacterial Blooms in Chilgok Weir, South Korea

Kim, J.; Jonoski, Andreja; Solomatine, Dmitri

doi:10.3390/w14040542

A Classification-Based Machine Learning Approach to the Prediction of Cyanobacterial Blooms in Chilgok Weir, South Korea

Journal article (2022)

Authors

J. Kim IHE Delft Institute for Water Education, Water Resources - , Human Resources Development Institute

Andreja Jonoski IHE Delft Institute for Water Education

Dmitri Solomatine IHE Delft Institute for Water Education, Water Resources - , Russian Academy of Sciences

Research Group

Water Resources () (TU Delft)

DOI: https://doi.org/10.3390/w14040542

Machine learning Feature selection Cyanobacterial blooms Classification algorithm Imbalanced dataset Oversampling

To reference this document use:

http://resolver.tudelft.nl/uuid:e11bdfc7-f0e5-4358-9410-c5e9f67eb54b

More Info

expand_more

Published Date

2022

Language

English

Faculty

Civil Engineering & Geosciences

Department

Water Management

Research Group

Water Resources

Abstract

Cyanobacterial blooms appear by complex causes such as water quality, climate, and hydrological factors. This study aims to present the machine learning models to predict occurrences of these complicated cyanobacterial blooms efficiently and effectively. The dataset was classified into groups consisting of two, three, or four classes based on cyanobacterial cell density after a week, which was used as the target variable. We developed 96 machine learning models for Chilgok weir using four classification algorithms: k-Nearest Neighbor, Decision Tree, Logistic Regression, and Support Vector Machine. In the modeling methodology, we first selected input features by applying ANOVA (Analysis of Variance) and solving a multi-collinearity problem as a process of feature selection, which is a method of removing irrelevant features to a target variable. Next, we adopted an oversampling method to resolve the problem of having an imbalanced dataset. Consequently, the best performance was achieved for models using datasets divided into two classes, with an accuracy of 80% or more. Comparatively, we confirmed low accuracy of approximately 60% for models using datasets divided into three classes. Moreover, while we produced models with overall high accuracy when using logCyano (logarithm of cyanobacterial cell density) as a feature, several models in combination with air temperature and NO3-N (nitrate nitrogen) using two classes also demonstrated more than 80% accuracy. It can be concluded that it is possible to develop very accurate classification-based machine learning models with two features related to cyanobacterial blooms. This proved that we could make efficient and effective models with a low number of inputs.

Files

Water_14_00542.pdf

(pdf | 29.3 Mb)