Exploiting Wikipedia as External Knowledge for Document Clustering

From Wikipedia Quality
Jump to: navigation, search


Exploiting Wikipedia as External Knowledge for Document Clustering
Authors
Xiaohua Hu
Xiaodan Zhang
Caimei Lu
E. K. Park
Xiaohua Zhou
Publication date
2009
DOI
10.1145/1557019.1557066
Links
Original

Exploiting Wikipedia as External Knowledge for Document Clustering - scientific work related to Wikipedia quality published in 2009, written by Xiaohua Hu, Xiaodan Zhang, Caimei Lu, E. K. Park and Xiaohua Zhou.

Overview

In traditional text clustering methods, documents are represented as "bags of words" without considering the semantic information of each document. For instance, if two documents use different collections of core words to represent the same topic, they may be falsely assigned to different clusters due to the lack of shared core words, although the core words they use are probably synonyms or semantically associated in other forms. The most common way to solve this problem is to enrich document representation with the background knowledge in an ontology. There are two major issues for this approach: (1) the coverage of the ontology is limited, even for WordNet or Mesh, (2) using ontology terms as replacement or additional features may cause information loss, or introduce noise. In this paper, authors present a novel text clustering method to address these two issues by enriching document representation with Wikipedia concept and category information. Authors develop two approaches, exact match and relatedness-match, to map text documents to Wikipedia concepts, and further to Wikipedia categories. Then the text documents are clustered based on a similarity metric which combines document content information, concept information as well as category information. The experimental results using the proposed clustering framework on three datasets (20-newsgroup, TDT2, and LA Times) show that clustering performance improves significantly by enriching document representation with Wikipedia concepts and categories.

Embed

Wikipedia Quality

Hu, Xiaohua; Zhang, Xiaodan; Lu, Caimei; Park, E. K.; Zhou, Xiaohua. (2009). "[[Exploiting Wikipedia as External Knowledge for Document Clustering]]".DOI: 10.1145/1557019.1557066.

English Wikipedia

{{cite journal |last1=Hu |first1=Xiaohua |last2=Zhang |first2=Xiaodan |last3=Lu |first3=Caimei |last4=Park |first4=E. K. |last5=Zhou |first5=Xiaohua |title=Exploiting Wikipedia as External Knowledge for Document Clustering |date=2009 |doi=10.1145/1557019.1557066 |url=https://wikipediaquality.com/wiki/Exploiting_Wikipedia_as_External_Knowledge_for_Document_Clustering}}

HTML

Hu, Xiaohua; Zhang, Xiaodan; Lu, Caimei; Park, E. K.; Zhou, Xiaohua. (2009). &quot;<a href="https://wikipediaquality.com/wiki/Exploiting_Wikipedia_as_External_Knowledge_for_Document_Clustering">Exploiting Wikipedia as External Knowledge for Document Clustering</a>&quot;.DOI: 10.1145/1557019.1557066.