sARI: a soft agreement measure for class partitions incorporating assignment probabilities

Flynt, A., Dean, N. and Nugent, R. (2019) sARI: a soft agreement measure for class partitions incorporating assignment probabilities. Advances in Data Analysis and Classification, 13(1), pp. 303-323. (doi:10.1007/s11634-018-0346-x)

[img]
Preview
Text
169741.pdf - Accepted Version

617kB

Abstract

Agreement indices are commonly used to summarize the performance of both classification and clustering methods. The easy interpretation/intuition and desirable properties that result from the Rand and adjusted Rand indices, has led to their popularity over other available indices. While more algorithmic clustering approaches like k-means and hierarchical clustering produce hard partition assignments (assigning observations to a single cluster), other techniques like model-based clustering include information about the certainty of allocation of objects through class membership probabilities (soft partitions). To assess performance using traditional indices, e.g., the adjusted Rand index (ARI), the soft partition is mapped to a hard set of assignments, which commonly overstates the certainty of correct assignments. This paper proposes an extension of the ARI, the soft adjusted Rand index (sARI), with similar intuition and interpretation but also incorporating information from one or two soft partitions. It can be used in conjunction with the ARI, comparing the similarities of hard to soft, or soft to soft partitions to the similarities of the mapped hard partitions. Simulation study results support the intuition that in general, mapping to hard partitions tends to increase the measure of similarity between partitions. In applications, the sARI more accurately reflects the cluster boundary overlap commonly seen in real data.

Item Type:Articles
Status:Published
Refereed:Yes
Glasgow Author(s) Enlighten ID:Dean, Dr Nema
Authors: Flynt, A., Dean, N., and Nugent, R.
College/School:College of Science and Engineering > School of Mathematics and Statistics > Statistics
Journal Name:Advances in Data Analysis and Classification
Publisher:Springer
ISSN:1862-5347
ISSN (Online):1862-5355
Published Online:09 October 2018
Copyright Holders:Copyright © 2018 Springer-Verlag GmbH Germany, part of Springer Nature
First Published:First published in Advances in Data Analysis and Classification 13:303-323
Publisher Policy:Reproduced in accordance with the copyright policy of the publisher

University Staff: Request a correction | Enlighten Editors: Update this record