Skip to Main Content (Press Enter)

Logo CNR
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze

UNI-FIND
Logo CNR

|

UNI-FIND

cnr.it
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze
  1. Pubblicazioni

Entity deduplication in big data graphs for scholarly communication

Articolo
Data di Pubblicazione:
2020
Abstract:
Purpose: Several online services offer functionalities to access information from "big research graphs" (e.g. Google Scholar, OpenAIRE, Microsoft Academic Graph), which correlate scholarly/scientific communication entities such as publications, authors, datasets, organizations, projects, funders, etc. Depending on the target users, access can vary from search and browse content to the consumption of statistics for monitoring and provision of feedback. Such graphs are populated over time as aggregations of multiple sources and therefore suffer from major entity-duplication problems. Although deduplication of graphs is a known and actual problem, existing solutions are dedicated to specific scenarios, operate on flat collections, local topology-drive challenges and cannot therefore be re-used in other contexts. Design/methodology/approach: This work presents GDup, an integrated, scalable, general-purpose system that can be customized to address deduplication over arbitrary large information graphs. The paper presents its high-level architecture, its implementation as a service used within the OpenAIRE infrastructure system and reports numbers of real-case experiments. Findings: GDup provides the functionalities required to deliver a fully-fledged entity deduplication workflow over a generic input graph. The system offers out-of-the-box Ground Truth management, acquisition of feedback from data curators and algorithms for identifying and merging duplicates, to obtain an output disambiguated graph. Originality/value: To our knowledge GDup is the only system in the literature that offers an integrated and general-purpose solution for the deduplication graphs, while targeting big data scalability issues. GDup is today one of the key modules of the OpenAIRE infrastructure production system, which monitors Open Science trends on behalf of the European Commission, National funders and institutions.
Tipologia CRIS:
01.01 Articolo in rivista
Keywords:
deduplication; information graphs; big data; scholarly communication; scalability; implementation
Elenco autori:
DE BONIS, Michele; Manghi, Paolo; Bardi, Alessia; Atzori, Claudio
Autori di Ateneo:
ATZORI CLAUDIO
BARDI ALESSIA
MANGHI PAOLO
Link alla scheda completa:
https://iris.cnr.it/handle/20.500.14243/385156
Link al Full Text:
https://iris.cnr.it//retrieve/handle/20.500.14243/385156/65379/prod_432254-doc_154523.pdf
Pubblicato in:
DATA TECHNOLOGIES AND APPLICATIONS
Journal
  • Dati Generali

Dati Generali

URL

https://www.emerald.com/insight/content/doi/10.1108/DTA-09-2019-0163/full/html
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.5.0.0 | Sorgente dati: PREPROD (Ribaltamento disabilitato)