Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus

From Wikipedia Quality
Jump to: navigation, search


Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus
Authors
Samuel Reese
Gemma Boleda
Montse Cuadros
Lluís Padró
German Rigau
Publication date
2010
Links
Original

Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus - scientific work related to Wikipedia quality published in 2010, written by Samuel Reese, Gemma Boleda, Montse Cuadros, Lluís Padró and German Rigau.

Overview

This article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the Wikipedia and has been automatically enriched with linguistic information. To knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the open source library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assigns WordNet senses, and WordNet has been aligned across languages via the InterLingual Index, this sort of annotation opens the way to massive explorations in lexical semantics that were not possible before. Authors present a first attempt at creating a trilingual lexical resource from the sense-tagged Wikipedia corpora, namely, WikiNet. Moreover, authors present two by-products of the project that are of use for the NLP community: An open source Java-based parser for Wikipedia pages developed for the construction of the corpus, and the integration of the WSD algorithm UKB in FreeLing.

Embed

Wikipedia Quality

Reese, Samuel; Boleda, Gemma; Cuadros, Montse; Padró, Lluís; Rigau, German. (2010). "[[Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus]]".

English Wikipedia

{{cite journal |last1=Reese |first1=Samuel |last2=Boleda |first2=Gemma |last3=Cuadros |first3=Montse |last4=Padró |first4=Lluís |last5=Rigau |first5=German |title=Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus |date=2010 |url=https://wikipediaquality.com/wiki/Wikicorpus:_a_Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus}}

HTML

Reese, Samuel; Boleda, Gemma; Cuadros, Montse; Padró, Lluís; Rigau, German. (2010). &quot;<a href="https://wikipediaquality.com/wiki/Wikicorpus:_a_Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus">Wikicorpus: a Word-Sense Disambiguated Multilingual Wikipedia Corpus</a>&quot;.