%0 Journal Article %9 ACL : Articles dans des revues avec comité de lecture répertoriées par l'AERES %A Serrano Valderas, Eva C. %A Berti-Equille, Laure %A Armienta Hernandez, M.A. %A Grac, C. %T Principled data preprocessing application biological aquatic indicators of water pollution %C Piscataway %D 2017 %L fdi:010073005 %G ENG %I IEEE %K INFORMATIQUE SCIENTIFIQUE ; SYSTEME EXPERT ; TRAITEMENT DE DONNEES ; ANALYSE STATISTIQUE ; QUALITE ; POLLUTION BIOLOGIQUE ; INDICATEUR ECOLOGIQUE %K BIOINFORMATIQUE ; FOUILLE DE DONNEES %M ISI:000426078300011 %P 5 %R 10.1109/DEXA.2017.27 %U https://www.documentation.ird.fr/hor/fdi:010073005 %> https://www.documentation.ird.fr/intranet/publi/depot/2018-07-10/010073005.pdf %W Horizon (IRD) %X In many biological studies, statistical and data mining methods are extensively used to analyze the data and discover actionable knowledge. But, bad data quality causing incorrect analysis results and wrong interpretations may induce misleading conclusions and inadequate decisions. To ensure the validity of the results, avoid bias and data misuse, it is necessary to control not only the whole analytical pipeline, but most importantly the quality of the data with appropriate data preprocessing choices. Since various preprocessing techniques and alternative strategies may lead to dramatically different outputs, it is crucial to rely on a principled and rigorous method to select the optimal set of data preprocessing steps that depends both on the input data distributional characteristics and on the inherent characteristics of the targeted statistical or data mining methods. In this paper, we propose a method that selects, given a dataset, the optimal set of preprocessing tasks to apply to the data such that the overall data preprocessing output maximizes the quality of the analytical results for various techniques of clustering, regression, and classification. We present some promising results that validate our approach on biomonitoring data preparation. %B DEXA : International Workshop on Database and Expert Systems Applications %8 2017/08/28-31 %$ 122APPLIC