Similarity Based Chinese Synonym Collocation Extraction

Li, Wanyin; Lu, Qin; Xu, Ruifeng

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Similarity Based Chinese Synonym Collocation Extraction
作者	Li, Wanyin (Li, Wanyin)、Lu, Qin (Lu, Qin)、Xu, Ruifeng (Xu, Ruifeng)
中文摘要	Collocation extraction systems based on pure statistical methods suffer from two major problems. The first problem is their relatively low precision and recall rates. The second problem is their difficulty in dealing with sparse collocations. In order to improve performance, both statistical and lexicographic approaches should be considered. This paper presents a new method to extract synonymous collocations using semantic information. The semantic information is obtained by calculating similarities from HowNet. We have successfully extracted synonymous collocations which normally cannot be extracted using lexical statistics. Our evaluation conducted on a 60MB tagged corpus shows that we can extract synonymous collocations that occur with very low frequency and that the improvement in the recall rate is close to 100%. In addition, compared with a collocation extraction system based on the Xtract system for English, our algorithm can improve the precision rate by about 44%.
起訖頁	123-143
關鍵詞	Lexical statistics、Synonymous collocations、Similarity、Semantic information
刊名	中文計算語言學期刊
期數	200503 (10:1期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	Aligning Parallel Bilingual Corpora Statistically with Punctuation Criteria