Difference between revisions of "Word-Sense Disambiguated Multilingual Wikipedia Corpus"
(+ Infobox work) |
(+ cat.) |
||
(One intermediate revision by one other user not shown) | |||
Line 9: | Line 9: | ||
== Overview == | == Overview == | ||
This article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the [[Wikipedia]] and has been automatically enriched with linguistic information. To knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the [[open source]] library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assigns[[WordNet]] senses, andWordNet has been aligned across languages via the InterLingual | This article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the [[Wikipedia]] and has been automatically enriched with linguistic information. To knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the [[open source]] library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assigns[[WordNet]] senses, andWordNet has been aligned across languages via the InterLingual | ||
+ | |||
+ | == Embed == | ||
+ | === Wikipedia Quality === | ||
+ | <code> | ||
+ | <nowiki> | ||
+ | Reese, Samuel; Torrent, Gemma Boleda; Oller, Montserrat Cuadros; Padró, Lluís; Claramunt, German Rigau. (2010). "[[Word-Sense Disambiguated Multilingual Wikipedia Corpus]]". | ||
+ | </nowiki> | ||
+ | </code> | ||
+ | |||
+ | === English Wikipedia === | ||
+ | <code> | ||
+ | <nowiki> | ||
+ | {{cite journal |last1=Reese |first1=Samuel |last2=Torrent |first2=Gemma Boleda |last3=Oller |first3=Montserrat Cuadros |last4=Padró |first4=Lluís |last5=Claramunt |first5=German Rigau |title=Word-Sense Disambiguated Multilingual Wikipedia Corpus |date=2010 |url=https://wikipediaquality.com/wiki/Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus}} | ||
+ | </nowiki> | ||
+ | </code> | ||
+ | |||
+ | === HTML === | ||
+ | <code> | ||
+ | <nowiki> | ||
+ | Reese, Samuel; Torrent, Gemma Boleda; Oller, Montserrat Cuadros; Padró, Lluís; Claramunt, German Rigau. (2010). &quot;<a href="https://wikipediaquality.com/wiki/Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus">Word-Sense Disambiguated Multilingual Wikipedia Corpus</a>&quot;. | ||
+ | </nowiki> | ||
+ | </code> | ||
+ | |||
+ | |||
+ | |||
+ | [[Category:Scientific works]] | ||
+ | [[Category:English Wikipedia]] | ||
+ | [[Category:Spanish Wikipedia]] | ||
+ | [[Category:Catalan Wikipedia]] |
Latest revision as of 08:30, 25 January 2021
Authors | Samuel Reese Gemma Boleda Torrent Montserrat Cuadros Oller Lluís Padró German Rigau Claramunt |
---|---|
Publication date | 2010 |
Links | Original |
Word-Sense Disambiguated Multilingual Wikipedia Corpus - scientific work related to Wikipedia quality published in 2010, written by Samuel Reese, Gemma Boleda Torrent, Montserrat Cuadros Oller, Lluís Padró and German Rigau Claramunt.
Overview
This article presents a new freely available trilingual corpus (Catalan, Spanish, English) that contains large portions of the Wikipedia and has been automatically enriched with linguistic information. To knowledge, this is the largest such corpus that is freely available to the community: In its present version, it contains over 750 million words. The corpora have been annotated with lemma and part of speech information using the open source library FreeLing. Also, they have been sense annotated with the state of the art Word Sense Disambiguation algorithm UKB. As UKB assignsWordNet senses, andWordNet has been aligned across languages via the InterLingual
Embed
Wikipedia Quality
Reese, Samuel; Torrent, Gemma Boleda; Oller, Montserrat Cuadros; Padró, Lluís; Claramunt, German Rigau. (2010). "[[Word-Sense Disambiguated Multilingual Wikipedia Corpus]]".
English Wikipedia
{{cite journal |last1=Reese |first1=Samuel |last2=Torrent |first2=Gemma Boleda |last3=Oller |first3=Montserrat Cuadros |last4=Padró |first4=Lluís |last5=Claramunt |first5=German Rigau |title=Word-Sense Disambiguated Multilingual Wikipedia Corpus |date=2010 |url=https://wikipediaquality.com/wiki/Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus}}
HTML
Reese, Samuel; Torrent, Gemma Boleda; Oller, Montserrat Cuadros; Padró, Lluís; Claramunt, German Rigau. (2010). "<a href="https://wikipediaquality.com/wiki/Word-Sense_Disambiguated_Multilingual_Wikipedia_Corpus">Word-Sense Disambiguated Multilingual Wikipedia Corpus</a>".