Skip to Main Content (Press Enter)

Logo CNR
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze

UNI-FIND
Logo CNR

|

UNI-FIND

cnr.it
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze
  1. Pubblicazioni

An approximate algorithm for maximum inner product search over streaming sparse vectors

Articolo
Data di Pubblicazione:
2023
Abstract:
Maximum Inner Product Search or top-k retrieval on sparse vectors is well-understood in information retrieval, with a number of mature algorithms that solve it exactly. However, all existing algorithms are tailored to text and frequency-based similarity measures. To achieve optimal memory footprint and query latency, they rely on the near stationarity of documents and on laws governing natural languages. We consider, instead, a setup in which collections are streaming--necessitating dynamic indexing--and where indexing and retrieval must work with arbitrarily distributed real-valued vectors. As we show, existing algorithms are no longer competitive in this setup, even against na"ive solutions. We investigate this gap and present a novel approximate solution, called Sinnamon, that can efficiently retrieve the top-k results for sparse real valued vectors drawn from arbitrary distributions. Notably, Sinnamon offers levers to trade-off memory consumption, latency, and accuracy, making the algorithm suitable for constrained applications and systems. We give theoretical results on the error introduced by the approximate nature of the algorithm, and present an empirical evaluation of its performance on two hardware platforms and synthetic and real-valued datasets. We conclude by laying out concrete directions for future research on this general top-k retrieval problem over sparse vectors.
Tipologia CRIS:
01.01 Articolo in rivista
Keywords:
Approximate Algorithms; Sparse Vectors; Maximum Inner Product Search
Elenco autori:
Nardini, FRANCO MARIA
Autori di Ateneo:
NARDINI FRANCO MARIA
Link alla scheda completa:
https://iris.cnr.it/handle/20.500.14243/438971
Link al Full Text:
https://iris.cnr.it//retrieve/handle/20.500.14243/438971/117172/prod_487716-doc_202714.pdf
Pubblicato in:
ACM TRANSACTIONS ON INFORMATION SYSTEMS
Journal
  • Dati Generali

Dati Generali

URL

https://dl.acm.org/doi/10.1145/3609797
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.5.0.0 | Sorgente dati: PREPROD (Ribaltamento disabilitato)