| 英文摘要 |
As large language models have increasingly demonstrated multimodal comprehension and semantic reasoning capabilities, the question of whether generative AI can reliably support teachers in scoring students' visual artifacts has become an important issue in educational assessment research. This study aimed to examine the scoring agreement between generative artificial intelligence (AI) and human raters in the assessment of STEM interdisciplinary flowcharts. Using 81 STEM interdisciplinary flowcharts produced by junior high school students as the research sample, this study compared the level of agreement between generative AI and human scoring. The results indicated that generative AI achieved moderate scoring agreement in structural and intermediate cognitive dimensions. However, lower agreement was observed in the higher-order cognitive dimension involving conceptual integration and cross-disciplinary reasoning. These findings suggest that, under clearly defined scoring rubrics and structured procedural criteria, generative AI may serve as a supplementary assessment tool rather than a full replacement for human judgment. The outcomes of this study provide empirical evidence for future STEM interdisciplinary assessment design and AI-assisted scoring applications. |