Skip to Main Content (Press Enter)

Logo CNR
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze

UNI-FIND
Logo CNR

|

UNI-FIND

cnr.it
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze
  1. Pubblicazioni

Sorting out the document identifier assignment problem

Articolo
Data di Pubblicazione:
2007
Abstract:
The compression of Inverted File indexes in Web Search Engines has received a lot of attention in these last years. Compressing the index not only reduces space occupancy but also improves the overall retrieval performance since it allows a better exploitation of the memory hierarchy. In this paper we are going to empirically show that in the case of collections of Web Documents we can enhance the performance of compression algorithms by simply assigning identifiers to documents according to the lexicographical ordering of the URLs. We will validate this assumption by comparing several assignment techniques and several compression algorithms on a quite large document collection composed by about six million documents. The results are very encouraging since we can improve the compression ratio up to 40% using an algorithm that takes about ninety seconds to finish using only 100 MB of main memory.
Tipologia CRIS:
01.01 Articolo in rivista
Keywords:
H.3 Information Storage and Retrieval; H.3.1 Content Analysis and Indexing. Indexing Methods; Identifier assignment; Indexing technique; Information retrieval index compression
Elenco autori:
Silvestri, Fabrizio
Link alla scheda completa:
https://iris.cnr.it/handle/20.500.14243/43612
  • Dati Generali

Dati Generali

URL

http://www.springerlink.com/content/y0755644n8n48627/
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.5.0.0 | Sorgente dati: PREPROD (Ribaltamento disabilitato)