Normal view MARC view ISBD view

Building a gold standard for event detection in Croatian / Ljubešić, Nikola ; Boras, Damir ; Lauc, Tomislava.

By: Ljubešić, Nikola, informatičar.
Contributor(s): Boras, Damir [aut] | Lauc, Tomislava [aut].
Material type: ArticleArticleDescription: 3101-3104 str.Other title: Building a gold standard for event detection in Croatian [Naslov na engleskom:].Subject(s): 5.04 | event detection, gold standard, newspaper text, Croatian language hrv | event detection, gold standard, newspaper text, Croatian language engOnline resources: Click here to access online In: Seventh International Conference on Language Resources and Evaluation (19.-21.05.2010. ; Valletta, Malta) Proceedings of the Seventh International Conference on Language Resources and Evaluation str. 3101-3104Calzolari, Nicoletta ; Choukri, Khalid ; Maegaard, Bente ; Mariani, Joseph ; Odjik, Jan ; Piperidis, Stelios ; Rosner, Mike ; Tapias, DanielSummary: This paper describes the process of building a newspaper corpus annotated with events described in specific documents. The main differ- ence to the corpora built as part of the TDT initiative is that documents are not annotated by topics, but by specific events they describe. Additionally, documents are gathered from sixteen sources and all documents in the corpus are annotated with the corresponding event. The annotation process consists of a browsing and a searching step. Experiments are performed with a threshold that could be used in the browsing step yielding the result of having to browse through only 1% of document pairs for a 2% loss of relevant document pairs. A statistical analysis of the annotated corpus is undertaken showing that most events are described by few documents while just some events are reported by many documents. The inter- annotator agreement measures show high agreement concerning grouping documents into event clusters, but show a much lower agreement concerning the number of events the documents are organized into. An initial experiment is described giving a baseline for further research on this corpus.
Tags from this library: No tags from this library for this title. Log in to add tags.
No physical items for this record

This paper describes the process of building a newspaper corpus annotated with events described in specific documents. The main differ- ence to the corpora built as part of the TDT initiative is that documents are not annotated by topics, but by specific events they describe. Additionally, documents are gathered from sixteen sources and all documents in the corpus are annotated with the corresponding event. The annotation process consists of a browsing and a searching step. Experiments are performed with a threshold that could be used in the browsing step yielding the result of having to browse through only 1% of document pairs for a 2% loss of relevant document pairs. A statistical analysis of the annotated corpus is undertaken showing that most events are described by few documents while just some events are reported by many documents. The inter- annotator agreement measures show high agreement concerning grouping documents into event clusters, but show a much lower agreement concerning the number of events the documents are organized into. An initial experiment is described giving a baseline for further research on this corpus.

Projekt MZOS 130-1301679-1380

ENG

There are no comments for this item.

Log in to your account to post a comment.

Powered by Koha

//