月旦知識庫
 
  1. 熱門:
 
首頁 臺灣期刊   法律   公行政治   醫事相關   財經   社會學   教育   其他 大陸期刊   核心   重要期刊 DOI文章
中文計算語言學期刊 本站僅提供期刊文獻檢索。
  【月旦知識庫】是否收錄該篇全文,敬請【登入】查詢為準。
最新【購點活動】


篇名
Modeling Taiwanese POS Tagging Using Statistical Methods and Mandarin Training Data
作者 Iunn, Un-gian (Iunn, Un-gian)Tai, Jia-hung (Tai, Jia-hung)Lau, Kiat-gak (Lau, Kiat-gak)Kao, Cheng-yan (Kao, Cheng-yan)Keh-Jiann Chen (Keh-Jiann Chen)
中文摘要
In this paper, we introduce a POS tagging method for Taiwan Southern Min. We use the more than 62,000 entries of the Taiwanese-Mandarin dictionary and 10 million words of Mandarin training data to tag Taiwanese. The literary written Taiwanese corpora have both Romanized script and Han-Romanization mixed script, and include prose, novels, and dramas. We follow the tagset drawn up by CKIP. We developed a word alignment checker to assist with the word alignment for the two scripts. It searches the Taiwanese-Mandarin dictionary to find corresponding Mandarin candidate words, selects the most suitable Mandarin word using an HMM probabilistic model from the Mandarin training data, and tags the word using an MEMM classifier. We achieve an accuracy rate of 91.6% on Taiwanese POS tagging work, and we analyze the errors. We also discover some preliminary Taiwanese training data.
起訖頁 237-256
關鍵詞 Taiwan southern minPOS taggingWritten TaiwaneseHidden morkov modelMaximal entropy markov model
刊名 中文計算語言學期刊  
期數 200909 (14:3期)
出版單位 中華民國計算語言學學會
該期刊-下一篇 A Thesaurus-Based Semantic Classification of English Collocations
 

新書閱讀



最新影音


優惠活動




讀者服務專線:+886-2-23756688 傳真:+886-2-23318496
地址:臺北市館前路28 號 7 樓 客服信箱
Copyright © 元照出版 All rights reserved. 版權所有,禁止轉貼節錄