15

2026-07

大语言模型在批判性思维倾向量表开发中的代理效度研究 —— 基于模拟样本与学生样本的对比分析

大语言模型在批判性思维倾向量表开发中的代理效度研究 —— 基于模拟样本与学生样本的对比分析 李鲁越    林文静    杜君磊    郑勤华[通讯作者] (北京师范大学 远程教育研究中心,北京 100875)   摘要:批判性思维倾向量表开发长期依赖大规模样本和高成本数据,而大语言模型(Large Language Model,LLM) 提供了模拟新路径,但其代理效度亟待验证。基于此,文章在基础指令与角色约束两种模式下,调用GLM-4-Plus、 GPT-4o 和 o1-preview 生成模拟样本,并与1477名中国八年级学生的真实样本对比,通过分析分布差异、信效度差异和算法偏见,结果显示:基础指令模式下模型的代理效度整体优于角色约束模式;GLM-4-Plus在响应数量和性别均衡方面表现突出;而o1-preview在基础指令模式下的模拟样本分布与内部一致性较好,GLM-4-Plus、GPT-4o在模型拟合上表现更优,但三者均存在性别偏见;所有模拟样本在“开放包容”维度分布偏离,且此维度的信度表现不佳,而“自我反思”维度的效度不佳。基于此结论,文章针对大语言模型的代理效度、模拟局限进行讨论,指出LLM可以有效模拟人类响应、辅助量表开发,但尚无法完全替代人类样本,且需警惕算法偏见。最后,文章从技术优化与量表开发两个方面提出大语言模型代理效度的提升路径。文章的研究有助于显著降低教育测量研究的成本、提升整体的研究效率与可及性,推动了批判性思维测评工具的智能化,并为人工智能辅助教育实证研究提供了新的方法论范式。   关键词:大语言模型;批判性思维;代理效度;模拟样本   点击查看原文:大语言模型在批判性思维倾向量表开发中的代理效度研究 —— 基于模拟样本与学生样本的对比分析  

14

2026-07

Computer- Assisted Performance- Based Assessment for Mental Health: A Scoping Review

Computer- Assisted Performance- Based Assessment for Mental Health: A Scoping Review Hanya Li   |   Shuang Li School of Educational Technology, Faculty of Education, Beijing Normal University, Beijing, People's Republic of China    ABSTRACT  Adolescent mental health is foundational to personal development, yet it faces escalating challenges globally. While traditional assessment methods lack objectivity and ecological validity, integrating computer- assisted technology (CAT) into performance- based assessments (PBAs) offers a promising pathway. This review, following the PRISMA- ScR reporting standard, analyzed 89 articles (2015–2025) to map the assessed components, CAT applications, and scenario diversity in mental health PBAs. Analysis revealed a research emphasis on mental disorders, with critical domains for adolescent development remaining significantly un derstudied. CATs significantly enhanced PBAs through data analysis, data acquisition, scenario creation, and tool digitization. PBA scenarios are diverse, demonstrating the adaptability of PBAs for multidimensional mental health assessment. Prioritizing the design of PBAs for social–emotional and adaptive assessment is critical for the early identification of adolescent mental health issues. Furthermore, advancing predictive analytics and leveraging large language models for feedback generation are promising ways to unlock CAT's potential in enhancing PBAs. Importantly, integrating and adapting scenarios from validated scales by CATs into PBAs could further enhance assessment typicality and reliability.   点击查看原文:Computer- Assisted Performance- Based Assessment for Mental Health: A Scoping Review  

14

2026-07

LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items

LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items Junlei Du ,    Yishen Song ,    and Qinhua Zheng   Abstract—Nonanchor equating presents a significant challenge in educational assessment when test forms lack common items, requiring innovative solutions to ensurescore comparability across different test administrations. This study proposes a novel large language model-simulated nonequivalent groups with anchor test (LLM-SNGAT) method that leverages large language models (LLMs)tosimulatetest-takingsamplesandgeneratecommonitem sets for equating purposes. The approach eliminates traditional dependencies onspecialized test design and extensive demographic data collection by utilizing the inherent capabilities of LLMs to simulate diverse response patterns. We evaluated the method using Tucker and Levine equating approaches across multiple LLMs,including generative pre-trained transformer 4o (GPT-4o), O1-preview, and DeepSeek-R1. Results demonstrated the feasibil ity of the proposed approach, with the Tucker method showing superior performance and consistent improvements as common item coverage increased. Sensitivity analysis confirmed that model performance rankings remained consistent across varying prompt formulations. The study revealed characteristic that standard er rors were smallest near the mean and became larger farther away from the mean, and identified optimal common item proportions of 30%–50%forstableequating performance. While current limi tations include the capacity of LLMs toaccurately simulate human cognitive and behavioral diversity, this proof-of-concept study pro vides preliminary evidence for the feasibility of the LLM-SNGAT methodology. The approach represents a paradigm shift from resource-intensive traditional methods to computationally driven solutions, offering promising prospects for addressing nonanchor equating challenges in the digital age.   Index Terms—Artificial intelligence (AI)-assisted assessment, educational measurement, large language models (LLMs), nonanchor equating, proof-of-concept, simulated samples, test equating.   点击查看原文:LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items  

14

2026-07

科学探究活动场景下基于多模态数据的学生坚毅力表现分析 —— 基于B大学附属学校四年级学生的调查

科学探究活动场景下基于多模态数据的学生坚毅力表现分析 —— 基于B大学附属学校四年级学生的调查 郭利明1     郑勤华2[通讯作者]     齐欣3 (1.温州大学 教育学院,浙江温州 325035;   2.北京师范大学 远程教育研究中心,北京 100875;   3.中国科学技术馆,北京 100101)   摘要:坚毅力作为学生综合素质的重要组成部分,能够在探究实践活动场景中得到培养与发展。然而,由于传统测评范式的局限性,学生在真实场景表现中的坚毅力发展现状尚未得到充分揭示。为此,文章基于B大学附属学校486名四年级学生参与科学探究活动的真实性多模态表现数据,采用总体分析、类型分析、差异分析和调节效应分析对学生坚毅力的表现进行了探究。研究发现,样本校四年级学生的坚毅力总体水平偏低,内部结构发展不均衡;学生的坚毅力类型呈现多元化趋势;学生坚毅力的差异主要由个体内在行为与情感因素驱动,环境背景的影响有限;学生坚毅力内部具有跨群体的普适性与稳定性。基于以上结论,文章从课程与任务设计、教师支持与家校协同、评估诊断与技术赋能三个层面提出对策建议,以期为未来学生坚毅力的培养与发展提供借鉴。   关键词:坚毅力测评;坚毅力指数;多模态数据;科学探究活动;表现性评价   点击查看原文:科学探究活动场景下基于多模态数据的学生坚毅力表现分析 —— 基于B大学附属学校四年级学生的调查  

13

2026-07

大语言模型在教育研究样本模拟中的保真度与偏差分析——以在线自我调节学习能力为例

大语言模型在教育研究样本模拟中的保真度与偏差分析 —— 以在线自我调节学习能力为例   杜君磊   李爽   陈靖茜   李晶   [摘 要]  利用大语言模型生成模拟样本开展实验或调研,已成为教育研究范式革新的重要方向。 然而,当前鲜有研究通过与真实样本对比,系统检验模拟样本在教育研究中的可行性。 为此,本研究聚焦混合学习情境中的在线自我调节学习能力调研,基于173名7—8年级真实学生的基本特征(基础型)与在线学习行为指标(增强型)设计了两类提示词,系统比较了GLM-4-plus、GPT-4o 与 o1-preview 三种大语言模型所生成的模拟样本在信效度、数据分布及假设检验等方面的表现。 研究结果显示,GPT-4o与o1-preview在信效度、性别分布、亚群特征,以及与在线行为指标的相关性上具有较高保真度,而GLM-4-plus的整体表现相对逊色。此外,增强型提示词有助于提升模拟样本的结构效度并缓解性别偏差。研究还发现,与人类样本相比,模拟样本在多样性与变量关系上存在一定误差。本研究为理解大 语言模型关于自我调节学习能力的“机器思维”提供了实证依据,也为在教育研究中应用大语言模型生成的模拟样本提供了有价值的参考。   [关键词]  大语言模型;教育研究方法;模拟样本;在线自我调节学习   点击查看原文:大语言模型在教育研究样本模拟中的保真度与偏差分析——以在线自我调节学习能力为例    

13

2026-07

基于场景多模态数据的学生坚毅力智能化测评研究

基于场景多模态数据的学生坚毅力智能化测评研究  郑勤华1,郭利明2①,齐欣3 1.北京师范大学 远程教育研究中心,北京100875       2.温州大学 教育学院,浙江 温州325035       3.中国科学技术馆,北京100101   摘要:学生坚毅力的重要性日益明显。目前对学生坚毅力的测评大部分是以主观方法、单一的静态性文本模态数据为 主,鲜有研究基于场景融合真实表现的多模态数据开展对学生坚毅力的客观化与智能化测评,不利于学生坚毅力的培养与发 展。基于此,该研究以学生坚毅力测评的理论体系、工具体系以及数据体系为基础,通过“多模态数据采集、多模态数据特征提取、模态融合与指标计算、模型生成与结果计算、模型有效性验证”的思路,开展融合科学探究活动场景多模态数据的学生坚毅力智能化测评。实践结果表明,该研究提出的学生坚毅力智能化测评方法体系能够比较真实反映学生的坚毅力发展 状态,一定程度解决了学生坚毅力测评的现实问题,能为未来学生高阶能力集的测评提供参考。   关键词:坚毅力测评;多模态数据;坚毅力测评可计算模型;表现性评价;科学探究活动   本文系国家自然科学基金面上项目“基于多模态数据融合计算的中小学生坚毅力测评技术与溯源研究”(项目编号:62277004)、 浙江省教育科学规划2025年一般规划课题(高校)“基于场景多模态数据的学生坚毅力测评与培养机制研究”(课题编号:2025SCG166) 阶段性研究成果。 ①郭利明为本文通讯作者   点击查看原文:基于场景多模态数据的学生坚毅力智能化测评研究  

31

2024-03

Automated Essay Scoring and Revising Based on Open-Source Large Language Models

Automated Essay Scoring and Revising Based onOpen-Source Large Language Models Yishen Song , Qianta Zhu , Huaibo Wang , and Qinhua Zheng Abstract—Manually scoring and revising student essays has long been a time-consuming task for educators.With the rise of natural language processing techniques, automated essay scoring (AES) and automated essay revising (AER) have emerged to alleviate this burden. However, current AES and AER models require large amounts of training data and lack generalizability, which makes them hard to implement in daily teaching activities. Moreover,online sites offering AES and AER services charge high fees and have security issues uploading student content. In light of these challenges and recognizing the advancements in large language models (LLMs), we aim to fill these research gaps by analyzing the performance of open-source LLMs when accomplishing AES and AER tasks. Using a human-scored essay dataset (n = 600) collected in an online assessment, we implemented zero-shot, few-shot, and p-tuning AES methods based on theLLMsand conducted a human–machine consistency check. We conducted a similarity test and a score difference test for the results of AER with LLMs support. The human–machine consistency check result shows that the performance of open-sourceLLMswith a 10Bparameter size in the AES task is close to that of some deep-learning baseline models,and it can be improved by integrating the comment with the score into the shot or training continuous prompts. The similarity test and score difference test results show that open-source LLMs can effectively accomplish the AER task, improving the quality of the essays while ensuring that the revision results are similar to the original essays. This study reveals a practical path to cost-effectively,time-efficiently, and content-safely assisting teachers with student essay scoring and revising using open-source LLMs. Index Terms—Assessment, automated essay revising (AER),automated essay scoring (AES), generative artificial intelligence,open-source large language model (LLM). 点击查看原文:Automated Essay Scoring and Revising Based on Open-Source Large Language Models