ISSN 0439-755X
CN 11-1911/B
主办:中国心理学会
   中国科学院心理研究所
出版:科学出版社

心理学报 ›› 2026, Vol. 58 ›› Issue (10): 2139-2152.doi: 10.3724/SP.J.1041.2026.2139 cstr: 32110.14.2026.2139

• 研究报告 • 上一篇    下一篇

项目反应理论下的残差加权稳健项目参数估计方法

高旭亮, 李婷婷, 王芳   

  1. 贵州师范大学心理学院, 贵州师范大学心理健康教育与咨询中心, 贵阳 550025
  • 收稿日期:2025-11-03 发布日期:2026-08-04 出版日期:2026-10-25
  • 通讯作者: 王芳, E-mail: wangf9817@foxmail.com
  • 基金资助:
    国家自然科学基金项目(32460212)、贵州省基础研究计划(自然科学)面上项目(黔科合基础MS[2025] 261)和2024年贵州省高校心理健康教育专项项目(JYT-XLZX-2024-BK019)资助

A residual-based weighted robust item parameter estimation method for aberrant responses

GAO Xuliang, LI Tingting, WANG Fang   

  1. School of Psychology Guizhou Normal University; Mental Health Education and Counseling Center, Guizhou Normal University, Guiyang 550025, China
  • Received:2025-11-03 Online:2026-08-04 Published:2026-10-25

摘要: 异常作答在心理与教育测评中普遍存在, 易导致项目参数估计偏差, 并削弱统计推断的准确性与解释力。为应对这一问题, 本研究在项目反应理论框架下提出了新的加权参数估计方法RRW, 并通过模拟实验与实证研究对其有效性进行了检验。实验结果显示, RRW方法在多种异常作答情境下均表现出较低的RMSE和Bias, 尤其在高异常比例条件下, 项目参数估计精度在部分情境下较传统方法和$ l^*_z$加权方法提升幅度可超过30%。实证分析进一步验证了RRW方法在真实测验数据中的优越性, 其能够在提升项目参数估计精度的同时, 提高正常作答者的能力估计准确性。综上, RRW加权估计在复杂异常作答情境下具有显著优势与良好应用前景。

关键词: 项目反应理论, 异常作答, 稳健项目参数估计

Abstract: Item Response Theory (IRT) models are widely used in psychological and educational measurement. Their effectiveness, however, can be undermined by data quality issues, particularly aberrant responding, such as cheating, careless errors, rapid guessing, and hybrid forms involving multiple aberrant response behaviors. Such responses are especially prevalent in online assessments, where the absence of supervision, external distractions, and low participant motivation can introduce substantial measurement noise. Empirical evidence indicates that even 10~15% proportion of aberrant responses can meaningfully distort parameter estimates, while higher proportions (20~30% or more) can lead to severe bias in item and ability parameter estimates, reduce reliability, distort factor structures, and compromise measurement invariance test.
Person-fit statistics (PFS) are widely used to identify aberrant respondents. However, their dependence on full-sample parameter estimates makes them susceptible to the masking effect, in which aberrant responses bias item parameter estimation and, in turn, obscure person misfit, particularly when the proportion of aberrant responses is high. To overcome this limitation, a robust marginal maximum likelihood method has been proposed that incorporates PFS-derived weights into the likelihood function, thereby down-weighting aberrant respondents during parameter estimation. Using the $ l^*_z$ statistic, this method can effectively reduce bias while retaining all response data.
Building on this framework, the present study introduces a new robust weighting method, namely Residual-Based Robust Weighting (RRW), which replaces $ l^*_z$ with a residual-based person-fit statistic (PFS). Residual-based statistics quantify item-level deviations between observed and expected responses, standardized by their variances and aggregated across items. This finer-grained approach improves sensitivity to subtle and heterogeneous aberrant patterns compared to the aggregated likelihood ratio underlying $ l^*_z$. As a result, RRW assigns lower weights to respondents whose residual patterns deviate substantially from model expectations, while preserving the full contribution of respondents whose responses are consistent with the model.
We evaluated the proposed RRW method through simulation studies and an empirical example, comparing its performance with unweighted estimation and the $ l^*_z$-based RMML approach. Results showed that aberrant responding substantially distorts item parameter estimation, with the unweighted method exhibiting the largest increases in RMSE and bias as aberrance levels rise. The $ l^*_z$-based method improved discrimination estimates to some extent but showed limited effectiveness for threshold parameters. In contrast, RRW consistently achieved lower estimation errors for both discrimination and threshold parameters across most conditions, with its advantages becoming more pronounced under higher aberrance levels. Moreover, RRW maintained stable performance as test length increased, demonstrating strong robustness and applicability in complex testing scenarios.
Empirical results further validated RRW's superiority: it reduced item parameter estimation error to a certain extent and enhanced ability estimates for non-aberrant respondents, demonstrating its effectiveness in mitigating the adverse impact of aberrant responses in real testing contexts. By improving item parameter accuracy, RRW indirectly increases the precision of ability estimation, providing a practical and robust alternative for IRT applications in complex testing environments. Overall, evidence from both simulation and empirical analyses indicates that the RRW weighted estimation method effectively reduces the negative effects of complex aberrant response patterns on item parameter estimation and shows strong potential for practical implementation.

Key words: item response theory, aberrant responding, robust item parameter estimation

中图分类号: