Difference between revisions of "Information Technology and Quantitative Management , Itqm 2014 a Method for Refining a Taxonomy by Using Annotated Suffix Trees and Wikipedia Resources"

From Wikipedia Quality
Jump to: navigation, search
(Starting an article - Information Technology and Quantitative Management , Itqm 2014 a Method for Refining a Taxonomy by Using Annotated Suffix Trees and Wikipedia Resources)
 
(Adding wikilinks)
Line 1: Line 1:
'''Information Technology and Quantitative Management , Itqm 2014 a Method for Refining a Taxonomy by Using Annotated Suffix Trees and Wikipedia Resources''' - scientific work related to Wikipedia quality published in 2014, written by Ekaterina Chernyak and Boris Mirkin.
+
'''Information Technology and Quantitative Management , Itqm 2014 a Method for Refining a Taxonomy by Using Annotated Suffix Trees and Wikipedia Resources''' - scientific work related to [[Wikipedia quality]] published in 2014, written by [[Ekaterina Chernyak]] and [[Boris Mirkin]].
  
 
== Overview ==
 
== Overview ==
A two-step approach to taxonomy construction is presented. On the first step the frame of taxonomy is built manually according to some representative educational materials. On the second step, the frame is refined using the Wikipedia category tree and articles. Since the structure of Wikipedia is rather noisy, a procedure to clear the Wikipedia category tree is suggested. A string-to-text relevance score, based on annotated suffix trees, is used several times to 1) clear the Wikipedia data from noise; 2) to assign W ikipedia categories to taxonomy topics; 3) to choose whether the category should be assigned to the taxonomy topic or stay on intermediate levels. The resulting taxonomy consists of three parts: the manully set upper levels, the adopted Wikipedia category tree and the Wikipedia articles as leaves.Also, a set of so-called descriptors is assigned to every leaf; these are phrases explaining aspects of the leaf topic. The method is illustrated by its application to two domains: a) Probability theory and mathematical statistics, b) ”Numerical analysis” (both in Russian). c
+
A two-step approach to taxonomy construction is presented. On the first step the frame of taxonomy is built manually according to some representative educational materials. On the second step, the frame is refined using the [[Wikipedia]] category tree and articles. Since the structure of Wikipedia is rather noisy, a procedure to clear the Wikipedia category tree is suggested. A string-to-text relevance score, based on annotated suffix trees, is used several times to 1) clear the Wikipedia data from noise; 2) to assign W ikipedia [[categories]] to taxonomy topics; 3) to choose whether the category should be assigned to the taxonomy topic or stay on intermediate levels. The resulting taxonomy consists of three parts: the manully set upper levels, the adopted Wikipedia category tree and the Wikipedia articles as leaves.Also, a set of so-called descriptors is assigned to every leaf; these are phrases explaining aspects of the leaf topic. The method is illustrated by its application to two domains: a) Probability theory and mathematical statistics, b) ”Numerical analysis” (both in Russian). c

Revision as of 23:09, 31 May 2019

Information Technology and Quantitative Management , Itqm 2014 a Method for Refining a Taxonomy by Using Annotated Suffix Trees and Wikipedia Resources - scientific work related to Wikipedia quality published in 2014, written by Ekaterina Chernyak and Boris Mirkin.

Overview

A two-step approach to taxonomy construction is presented. On the first step the frame of taxonomy is built manually according to some representative educational materials. On the second step, the frame is refined using the Wikipedia category tree and articles. Since the structure of Wikipedia is rather noisy, a procedure to clear the Wikipedia category tree is suggested. A string-to-text relevance score, based on annotated suffix trees, is used several times to 1) clear the Wikipedia data from noise; 2) to assign W ikipedia categories to taxonomy topics; 3) to choose whether the category should be assigned to the taxonomy topic or stay on intermediate levels. The resulting taxonomy consists of three parts: the manully set upper levels, the adopted Wikipedia category tree and the Wikipedia articles as leaves.Also, a set of so-called descriptors is assigned to every leaf; these are phrases explaining aspects of the leaf topic. The method is illustrated by its application to two domains: a) Probability theory and mathematical statistics, b) ”Numerical analysis” (both in Russian). c