Skip to Main Content (Press Enter)

Logo CNR
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze

UNI-FIND
Logo CNR

|

UNI-FIND

cnr.it
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze
  1. Pubblicazioni

Cascaded transformer-based networks for Wikipedia large-scale image-caption matching

Articolo
Data di Pubblicazione:
2024
Abstract:
With the increasing importance of multimedia and multilingual data in online encyclopedias, novel methods are needed to fill domain gaps and automatically connect different modalities for increased accessibility. For example,Wikipedia is composed of millions of pages written in multiple languages. Images, when present, often lack textual context, thus remaining conceptually floating and harder to find and manage. In this work, we tackle the novel task of associating images from Wikipedia pages with the correct caption among a large pool of available ones written in multiple languages, as required by the image-caption matching Kaggle challenge organized by theWikimedia Foundation.Asystem able to perform this task would improve the accessibility and completeness of the underlying multi-modal knowledge graph in online encyclopedias. We propose a cascade of two models powered by the recent Transformer networks able to efficiently and effectively infer a relevance score between the query image data and the captions. We verify through extensive experiments that the proposed cascaded approach effectively handles a large pool of images and captions while maintaining bounded the overall computational complexity at inference time.With respect to other approaches in the challenge leaderboard,we can achieve remarkable improvements over the previous proposals (+8% in nDCG@5 with respect to the sixth position) with constrained resources. The code is publicly available at https://tinyurl.com/wiki-imcap.
Tipologia CRIS:
01.01 Articolo in rivista
Keywords:
Multi-modal matching; Information retrieval; Deep learning; Transformer networks
Elenco autori:
Coccomini, DAVIDE ALESSANDRO; Falchi, Fabrizio; Esuli, Andrea; Messina, Nicola
Autori di Ateneo:
ESULI ANDREA
FALCHI FABRIZIO
Link alla scheda completa:
https://iris.cnr.it/handle/20.500.14243/453529
Link al Full Text:
https://iris.cnr.it//retrieve/handle/20.500.14243/453529/156447/prod_491916-doc_205202.pdf
Pubblicato in:
MULTIMEDIA TOOLS AND APPLICATIONS
Journal
  • Dati Generali

Dati Generali

URL

https://link.springer.com/article/10.1007/s11042-023-17977-0
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.5.0.0 | Sorgente dati: PREPROD (Ribaltamento disabilitato)