Data di Pubblicazione:
2016
Abstract:
The process of discovering interesting patterns in large, possibly huge, data sets is referred to as data mining, and can be performed in several flavours, known as "data mining functions". Among these functions, Outlier Detection discovers observations which deviate substantially from the rest of the data, and has many important practical applications. Outlier detection in very large data sets is however computationally very demanding and currently requires high-performance computing facilities. We propose a family of parallel and distributed algorithms for Graphic Processing Units (GPU) derived from two distance-based outlier detection algorithms: the BruteForce and the SolvingSet. The algorithms differ in the way they exploit the architecture and memory hierarchy of the GPU and guarantee significant improvements with respect to the CPU versions, both in terms of scalability and exploitation of parallelism. We provide a detailed discussion of their computational properties and measure performances with an extensive experimentation, comparing the several implementations and showing significant speedups.
Tipologia CRIS:
01.01 Articolo in rivista
Keywords:
Distance-based outliers; outlier detection; parallel and distributed algorithms
Elenco autori:
Basta, Stefano
Link alla scheda completa:
Pubblicato in: