月旦知識庫
 
  1. 熱門:
 
首頁 臺灣期刊   法律   公行政治   醫事相關   財經   社會學   教育   其他 大陸期刊   核心   重要期刊 DOI文章
中文計算語言學期刊 本站僅提供期刊文獻檢索。
  【月旦知識庫】是否收錄該篇全文,敬請【登入】查詢為準。
最新【購點活動】


篇名
Latent Semantic Language Modeling and Smoothing of Chinese Language
作者 Chien, Jen-tzung (Chien, Jen-tzung)Wu, Meng-sung (Wu, Meng-sung)Peng,Hua-jui (Peng,Hua-jui)
中文摘要
Language modeling plays a critical role for automatic speech recognition. Typically, the n-gram language models suffer from the lack of a good representation of historical words and an inability to estimate unseen parameters due to insufficient training data. In this study, we explore the application of latent semantic information (LSI) to language modeling and parameter smoothing. Our approach adopts latent semantic analysis to transform all words and documents into a common semantic space. The word-to-word, word-to-document and document-to-document relations are, accordingly, exploited for language modeling and smoothing. For language modeling, we present a new representation of historical words based on retrieval of the most relevant document. We also develop a novel parameter smoothing method, where the language models of seen and unseen words are estimated by interpolating the k nearest seen words in the training corpus. The interpolation coefficients are determined according to the closeness of words in the semantic space. As shown by experiments, the proposed modeling and smoothing methods can significantly reduce the perplexity of language models with moderate computational cost.
起訖頁 29-44
關鍵詞 Language modelingParameter smoothingSpeech recognitionLatent semantic analysis
刊名 中文計算語言學期刊  
期數 200408 (9:2期)
出版單位 中華民國計算語言學學會
該期刊-上一篇 Multiple-Translation Spotting for Mandarin-Taiwanese Speech-to-Speech Translation
該期刊-下一篇 Multi-Modal Emotion Recognition from Speech and Text
 

新書閱讀



最新影音


優惠活動




讀者服務專線:+886-2-23756688 傳真:+886-2-23318496
地址:臺北市館前路28 號 7 樓 客服信箱
Copyright © 元照出版 All rights reserved. 版權所有,禁止轉貼節錄