月旦知識庫
月旦知識庫 會員登入元照網路書店月旦品評家
 
 
  1. 熱門:
首頁 臺灣期刊   法律   公行政治   醫事相關   財經   社會學   教育   其他 大陸期刊   核心   重要期刊 DOI文章
ROCLING論文集 本站僅提供期刊文獻檢索。
  【月旦知識庫】是否收錄該篇全文,敬請【登入】查詢為準。
最新【購點活動】


篇名
Analyzing ChatGPT’s Mathematical Deficiencies: Insights and Contributions
並列篇名
Analyzing ChatGPT’s Mathematical Deficiencies: Insights and Contributions
英文摘要
In this study, we assess ChatGPT, OpenAI’s latest conversational chatbot and large language model (LLM), on its performance in elementary-grade arithmetic and logic problems. Despite its impressive coherence in natural language processing and ability to follow instructions, our findings indicate that Chat- GPT still has room for improvement in mathematical tasks. To evaluate its performance, we used six math and logic datasets, including SingleEq, AddSub, SVAMP, MultiArith, Simple Arithmetic and counting, and Arithmetic (word variation), and found that ChatGPT performed better than previous models such as InstructGPT and Minerva. However, our arithmetic dataset, which includes two- to seven-digit equations, revealed that ChatGPT’s accuracy in solving addition problems decreased from 100% to 64%, with simple arithmetic errors such as not carrying over in addition being a common issue. Additionally, the model struggled with basic multi-step word problems. To address this, we propose a novel benchmark for evaluating LLMs’mathematical abilities. Further research is needed for LLMs to reach the level of mathematical reasoning comparable to their natural language processing abilities. Overall, our study highlights the need for continued improvement in LLMs’mathematical abilities to make them more effective in real-world applications.
起訖頁 188-193
關鍵詞 Large language modelsreasoning capabilities
刊名 ROCLING論文集  
期數 202310 (2023期)
出版單位 中華民國計算語言學學會
該期刊-上一篇 Phonotactic Constraints on Zhangzhou Onsets
該期刊-下一篇 以語料庫及華語教材為本之「X是Y」隱喻探析
 

新書閱讀



最新影音


優惠活動




讀者服務專線:+886-2-23756688 傳真:+886-2-23318496
地址:臺北市館前路28 號 7 樓 客服信箱
Copyright © 元照出版 All rights reserved. 版權所有,禁止轉貼節錄