Skip to Main Content (Press Enter)

Logo CNR
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze

UNI-FIND
Logo CNR

|

UNI-FIND

cnr.it
  • ×
  • Home
  • Persone
  • Pubblicazioni
  • Strutture
  • Competenze
  1. Pubblicazioni

A utility-theoretic ranking method for semi-automated text classification.

Contributo in Atti di convegno
Data di Pubblicazione:
2012
Abstract:
In Semi-Automated Text Classification (SATC) an automatic classifier Phi labels a set of unlabelled documents D, following which a human annotator inspects (and corrects when appropriate) the labels attributed by Phi to a subset D' of D, with the aim of improving the overall quality of the labelling. An automated system can support this process by ranking the automatically labelled documents in a way that maximizes the expected increase in effectiveness that derives from inspecting D'. An obvious strategy is to rank D so that the documents that Phi has classified with the lowest confidence are top-ranked. In this work we show that this strategy is suboptimal. We develop a new utility-theoretic ranking method based on the notion of inspection gain, defined as the improvement in classification effectiveness that would derive by inspecting and correcting a given automatically labelled document. We also propose a new effectiveness measure for SATC-oriented ranking methods, based on the expected reduction in classification error brought about by partially inspecting a list generated by a given ranking method. We report the results of experiments showing that, with respect to the baseline method above, and according to the proposed measure, our ranking method can achieve substantially higher expected reductions in classification error.
Tipologia CRIS:
04.01 Contributo in Atti di convegno
Keywords:
cost-sensitive learning; ranking; semi-automated text classification; supervised learning; text classification
Elenco autori:
Esuli, Andrea; Berardi, Giacomo; Sebastiani, Fabrizio
Autori di Ateneo:
ESULI ANDREA
SEBASTIANI FABRIZIO
Link alla scheda completa:
https://iris.cnr.it/handle/20.500.14243/2684
  • Utilizzo dei cookie

Realizzato con VIVO | Designed by Cineca | 26.5.0.0 | Sorgente dati: PREPROD (Ribaltamento disabilitato)