Aligning More Words with High Precision for Small Bilingual Corpora

Ker, Sue J.; Chang, Jason S.

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Aligning More Words with High Precision for Small Bilingual Corpora
作者	Ker, Sue J. (Ker, Sue J.)、Chang, Jason S. (Chang, Jason S.)
中文摘要	In this paper, we propose an algorithm for identifying each word with its translations in a sentence and translation pair. Previously proposed methods require enormous amounts of bilingual data to train statistical word-by-word translation models. By taking a word-based approach, these methods align frequent words with consistent translations at a high precision rate. However, less frequent words or words with diverse translations generally do not have statistically significant evidence for confident alignment. Consequently, incomplete or incorrect alignments occur. Here, we attempt to improve on the coverage using class-based rules. An automatic procedure for acquiring such rules is also described. Experimental results confirm that the algorithm can align over 85% of word pairs while maintaining a comparably high precision rate, even when a small corpus is used in training.
起訖頁	63-95
關鍵詞	Word alignment、Machine readable dictionary and thesaurus、Bilingual corpus、Word sense disambiguation
刊名	中文計算語言學期刊
期數	199708 (2:2期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	Segmentation Standard for Chinese Natural Language Processing
該期刊-下一篇	An Unsupervised Iterative Method for Chinese New Lexicon Extraction