Noisy Channel Models for Corrupted Chinese Text Restoration and GB-to-Big5 Conversion

Chang, Chao-huang

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Noisy Channel Models for Corrupted Chinese Text Restoration and GB-to-Big5 Conversion
作者	Chang, Chao-huang (Chang, Chao-huang)
中文摘要	In this article, we propose a noisy channel/information restoration model for error recovery problems in Chinese natural language processing. A language processing system is considered as an information restoration process executed through a noisy channel. By feeding a large-scale standard corpus C into a simulated noisy channel, we can obtain a noisy version of the corpus N. Using N as the input to the language processing system (i.e., the information restoration process), we can obtain the output results C'. After that, the automatic evaluationmodule compares the original corpus C and the output results C', and computes the performance index (i.e., accuracy) automatically. The proposed model has been applied to two common and important problems related to Chinese NLP for the Internet: corrupted Chinese text restoration and GB-to-BIG5 conversion. Sinica Corpora version 1.0 and 2.0 are used in the experiment. The results show that the proposed model is useful and practical.
起訖頁	79-91
刊名	中文計算語言學期刊
期數	199808 (3:2期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	An Assessment of Character-based Chinese News Filtering Using Latent Semantic Indexing
該期刊-下一篇	Statistical Analysis of Mandarin Acoustic Units and Automatic Extraction of Phonetically Rich Sentences Based Upon a Very Large Chinese Text Corpus