Design and Evaluation of Approaches to Automatic Chinese Text Categorization

Tsay, Jyh-jong; Wang, Jing-doo

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Design and Evaluation of Approaches to Automatic Chinese Text Categorization
作者	Tsay, Jyh-jong (Tsay, Jyh-jong)、Wang, Jing-doo (Wang, Jing-doo)
中文摘要	In this paper, we propose and evaluate approaches to categorizing Chinese texts, which consist of term extraction, term selection, term clustering and text classification. We propose a scalable approach which uses frequency counts to identify left and right boundaries of possibly significant terms. We used the combination of term selection and term clustering to reduce the dimension of the vector space to a practical level. While the huge number of possible Chinese terms makes most of the machine learning algorithms impractical, results obtained in an experiment on a CAN news collection show that the dimension could be dramatically reduced to 1200 while approximately the same level of classification accuracy was maintained using our approach. We also studied and compared the performance of three well known classifiers, the Rocchio linear classifier, naive Bayes probabilistic classifier and k-nearest neighbors(kNN) classifier, when they were applied to categorize Chinese texts. Overall, kNN achieved the best accuracy, about 78.3%, but required large amounts of computation time and memory when used to classify new texts. Rocchio was very time and memory efficient, and achieved a high level of accuracy, about 75.4%. In practical implementation, Rocchio may be a good choice.
起訖頁	43-58
關鍵詞	詞彙選擇、文件分類、Term Clustering、Term Selection、Text Categorization
刊名	中文計算語言學期刊
期數	200008 (5:2期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	Adaptive Word Sense Disambiguation Using Lexical Knowledge in a Machine-readable Dictionary
該期刊-下一篇	Japanese-Chinese Cross-Language Information Retrieval: An Interlingua Approach