Unknown Word Detection for Chinese by a Corpus-based Learning Method

Keh-Jiann Chen; Bai, Ming-hong

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Unknown Word Detection for Chinese by a Corpus-based Learning Method
作者	Keh-Jiann Chen (Keh-Jiann Chen)、Bai, Ming-hong (Bai, Ming-hong)
中文摘要	One of the most prominent problems in computer processing of the Chinese language is identification of the words in a sentence. Since there are no blanks to mark word boundaries, identifying words is difficult because of segmentation ambiguities and occurrences of out-of-vocabulary words (i.e., unknown words). In this paper, a corpus-based learning method is proposedwhich derives sets of syntactic rules that are applied to distinguish monosyllabic words from monosyllabic morphemes which may be parts of unknown words or typographical errors. The corpus-based learning approach has the advantages of: 1. automatic rule learning, 2. automatic evaluation of the performance of each rule, and 3. balancing of recall and precision rates through dynamic rule set selection. The experimental results show that the rule set derived using the proposed method outperformed hand-crafted rules produced by human experts in detecting unknown words.
起訖頁	27-44
刊名	中文計算語言學期刊
期數	199802 (3:1期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	Analyzing the Performance of Message Understanding Systems
該期刊-下一篇	Meaning Representation and Meaning Instantiation for Chinese Nominals