| 英文摘要 |
Generative artificial intelligence and large language models (LLMs) are rapidly entering psychological research workflows, with broad applications in research writing and programming assistance, stimulus generation, annotation and analysis of natural language data, and quasi-experimental uses that simulate human responses and generate synthetic respondent data. Undeniably, LLMs substantially enhance research efficiency and overcome the instability resulting from fatigue and standard drift common in traditional manual coding, delivering highly consistent classification outputs when appropriately configured. Yet this article argues that LLMs are not merely efficiency tools but substantively reshape the inferential conditions of data processing, measurement construction, and statistical inference in psychological research. The paper first examines how the probabilistic nature of LLMs systematically produces bias, and subsequently outlines four methodological roles that LLMs play in psychological research. Drawing on empirical studies from clinical, cognitive, and social psychology, it analyzes how LLMs acting as measurement instruments or proxy annotators pose challenges to validity argumentation and error structures. To prevent misplacement of inferential levels, a five-stage checking procedure is proposed, together with practically operable inferential thresholds and operational guidelines, and a call is issued for the scholarly ecosystem to establish corresponding review mechanisms. While psychology must adopt new statistical correction techniques to confront the new sources of variance that LLMs introduce, the ultimate purpose remains to uphold the validity argumentation and error control traditions long developed within psychometrics. This amounts to a principle of“innovation in methods, continuity of standards,”which maintains the auditability and cumulativeness of psychological knowledge. |