Publications des scientifiques de l'IRD

Paradis Emmanuel. (2022). Reduced multidimensional scaling. Computational Statistics, 37, 91-105. ISSN 0943-4062.

Titre du document
Reduced multidimensional scaling
Année de publication
2022
Type de document
Article référencé dans le Web of Science WOS:000658117600001
Auteurs
Paradis Emmanuel
Source
Computational Statistics, 2022, 37, 91-105 ISSN 0943-4062
Dimension reduction is a common problem when analysing large data sets. The present paper proposes a method called reduced multidimensional scaling based on performing an initial standard multidimensional scaling on a reduced data set. This method faces the problem of finding a representative reduced sample. An algorithm is presented to perform this selection based on alternating sampling in outlier areas and observations in high density areas. A space is then constructed with the selected reduced sample by standard multidimentional scaling using pairwise distances. The observations not included in the reduced sample are then projected on the constructed space using Gower's formula in order to obtain a final representation of the whole data set. The only requirement is the ability to compute distances among observations. A simulation study showed that the proposed algorithm results performs well to detect outliers. Evaluation of running times suggests that the proposed method could run in a few hours with data sets that would take more than one year to analyse with standard multidimensional scaling. An application is presented with a dataset of 9547 DNA sequences of human immunodeficiency viruses.
Plan de classement
Sciences fondamentales / Techniques d'analyse et de recherche [020]
Localisation
Fonds IRD [F B010082101]
Identifiant IRD
fdi:010082101
Contact