月旦知識庫
 
  1. 熱門:
 
首頁 臺灣期刊   法律   公行政治   醫事相關   財經   社會學   教育   其他 大陸期刊   核心   重要期刊 DOI文章
ROCLING論文集 本站僅提供期刊文獻檢索。
  【月旦知識庫】是否收錄該篇全文,敬請【登入】查詢為準。
最新【購點活動】


篇名
Unknown Word Detection for Chinese by a Corpus-based Learning Method
作者 Keh-Jiann ChenMing-Hong Bai (Ming-Hong Bai)
英文摘要
One of the most prominent problems in computer processing of Chinese language is identification of the words in a sentence. Since there are no blanks to mark word boundaries, identifying words is difficult because of segmentation ambiguities and occurrences of out-of-vocabulary words (i.e. unknown words). In this paper, a corpus-based learning method is proposed which derives sets of syntactic rules that are applied to distinguish monosyllabic words from monosyllabic morphemes which may be parts of unknown words or typographical errors. The corpus-based learning approach has the advantages of 1. automatic rule learning, 2. automatic evaluation of the performance of each rule, and 3. balancing of recall and precision rates through dynamic rule set selection. The experimental results show that the rule set derived by the proposed method outperformed hand-crafted rules produced by human experts in detecting unknown words.
起訖頁 159-174
刊名 ROCLING論文集  
期數 1997 (1997期)
出版單位 國立高雄師範大學輔導與諮商研究所
該期刊-上一篇 Proper Name Extraction from Web Pages for Finding People in Internet
該期刊-下一篇 Analyzing the Complexity of a Domain With Respect To An Information Extraction Task
 

新書閱讀



最新影音


優惠活動




讀者服務專線:+886-2-23756688 傳真:+886-2-23318496
地址:臺北市館前路28 號 7 樓 客服信箱
Copyright © 元照出版 All rights reserved. 版權所有,禁止轉貼節錄