Construction and Automatization of a Minnan Child Speech Corpus with some Research Findings

Jane S. Tsay

月旦知識庫會員登入｜元照網路書店｜月旦品評家

熱門：

首頁

臺灣期刊 法律公行政治醫事相關財經社會學教育其他

大陸期刊 核心重要期刊

DOI文章

	本站僅提供期刊文獻檢索。　　【月旦知識庫】是否收錄該篇全文，敬請【登入】查詢為準。最新【購點活動】
篇名	Construction and Automatization of a Minnan Child Speech Corpus with some Research Findings
作者	Jane S. Tsay (Jane S. Tsay)
中文摘要	Taiwanese Child Language Corpus (TAICORP) is a corpus based on spontaneous conversations between young children and their adult caretakers in Minnan (Taiwan Southern Min) speaking families in Chiayi County, Taiwan. This corpus is special in several ways: (1) It is a Minnan corpus; (2) It is a speech-based corpus; (3) It is a corpus of a language that does not yet have a conventionalized orthography; (4) It is a collection of longitudinal child language data; (5) It is one of the largest child corpora in the world with about two million syllables in 497,426 lines (utterances) based on about 330 hours of recordings. Regarding the format, TAICORP adopted the Child Language Data Exchange System (CHILDES) [MacWhinney and Snow 1985; MacWhinney 1995] for transcribing and coding the recordings into machine-readable text. The goals of this paper are to introduce the construction of this speech-based corpus and at the same time to discuss some problems and challenges encountered. The development of an automatic word segmentation program with a spell-checker is also discussed. Finally, some findings in syllable distribution are reported.
起訖頁	411-441
關鍵詞	Minnan、Taiwan Southern Min、Taiwanese、Speech Corpus、Child Language、CHILDES、Automatic Word Segmentation
刊名	中文計算語言學期刊
期數	200712 (12:4期)
出版單位	中華民國計算語言學學會
該期刊-上一篇	Some Studies on Min-Nan Speech Processing
該期刊-下一篇	Automatic Pronunciation Assessment for Mandarin Chinese: Approaches and System Overview