A Multivariate Gaussian Mixture Model for Automatic Compound Word Extraction

Jing-Shin Chang; Keh-Yih Su

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	A Multivariate Gaussian Mixture Model for Automatic Compound Word Extraction
作者	Jing-Shin Chang (Jing-Shin Chang)、Keh-Yih Su (Keh-Yih Su)
英文摘要	An improved statistical model is proposed in this paper for extracting compound words from a text corpus. Traditional terminology extraction methods rely heavily on simple filtering-and-thresholding methods, which are unable to minimize the error counts objectively. Therefore, a method for minimizing the error counts is very desirable. In this paper, an improved statistical model is developed to integrate parts of speech information as well as other frequently used word association metrics to jointly optimize the extraction tasks. The features are modelled with a multivariate Gaussian mixture for handling the inter-feature correlations properly. With a training (resp. testing) corpus of 20715 (resp. 2301) sentences, the weighted precision & recall (WPR) can achieve about 84% for bigram compounds, and 86% for trigram compounds. The F-measure performances are about 82% for bigrams and 84% for trigrams.
起訖頁	123-142
刊名	ROCLING論文集
期數	1997 (1997期)
出版單位	國立高雄師範大學輔導與諮商研究所
該期刊-上一篇	THE APPLICATION OF THE SIMILARITIES BETWEEN THE MORPHEMES OF THE ENGLISH AND CHINESE LANGUAGES TO REPRESENT CHINESE CHARACTERS PHONETICALLY WITH ENGLISH LETTERS TO FACILITATE COMPUTER APPLICATIONS MANUALLY AND BY VOICE WITH THE CHRACTER-BASED LANGUAGES CHINESE, JAPANESE AND KOREAN
該期刊-下一篇	Proper Name Extraction from Web Pages for Finding People in Internet

新書閱讀

元照讀書館

優惠活動

月旦品評家

元照讀書館

．研討會新訊

月旦知識庫

月旦法律分析庫
月旦醫事法網
月旦會計財稅網

期刊數位服務

社群平台

讀者服務

關於元照

讀者服務專線：+886-2-23756688　傳真：+886-2-23318496
地址：臺北市館前路28 號 7 樓　客服信箱