LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items
LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items Junlei Du , Yishen Song , and Qinhua Zheng Abstract—Nonanchor equating presents a significant challenge in educational assessment when test forms lack common items, requiring innovative solutions to ensurescore comparability across different test administrations. This study proposes a novel large language model-simulated nonequivalent groups with anchor test (LLM-SNGAT) method that leverages large language models (LLMs)tosimulatetest-takingsamplesandgeneratecommonitem sets for equating purposes. The approach eliminates traditional dependencies onspecialized test design and extensive demographic data collection by utilizing the inherent capabilities of LLMs to simulate diverse response patterns. We evaluated the method using Tucker and Levine equating approaches across multiple LLMs,including generative pre-trained transformer 4o (GPT-4o), O1-preview, and DeepSeek-R1. Results demonstrated the feasibil ity of the proposed approach, with the Tucker method showing superior performance and consistent improvements as common item coverage increased. Sensitivity analysis confirmed that model performance rankings remained consistent across varying prompt formulations. The study revealed characteristic that standard er rors were smallest near the mean and became larger farther away from the mean, and identified optimal common item proportions of 30%–50%forstableequating performance. While current limi tations include the capacity of LLMs toaccurately simulate human cognitive and behavioral diversity, this proof-of-concept study pro vides preliminary evidence for the feasibility of the LLM-SNGAT methodology. The approach represents a paradigm shift from resource-intensive traditional methods to computationally driven solutions, offering promising prospects for addressing nonanchor equating challenges in the digital age. Index Terms—Artificial intelligence (AI)-assisted assessment, educational measurement, large language models (LLMs), nonanchor equating, proof-of-concept, simulated samples, test equating. 点击查看原文:LLM-Simulated Nonequivalent Groups With Anchor Test: A Novel Approach for Test Equating in the Absence of Traditional Anchor Items