| 英文摘要 |
Financial restatements have significant implications for financial statement users, the capital market, and the companies involved. This study employs machine learning to classify restatements. In this research, two datasets are used to decrease the imbalance of data types caused by each restatement classifications. Five algorithms were applied to predict the classifications, and the K-nearest neighbor (KNN) algorithm demonstrated superior performance across the two datasets, achieving 81.99% and 78.29%, respectively. Utilizing feature selection to enhance prediction, eXtreme Gradient Boosting (XGBoost) increased prediction accuracy by 6.48% compared to KNN, achieving 83.90%. It is also worth noting that ReliefF proved more effective than information gain in improving accuracy. In addition, feature selection identified nine risk variables for restatement, which were then compared with risk variables from previous empirical accounting research. |