Accelerate Literature Icon
Want to do a literature review? Try our new Literature Review workflow

Leveraging Process Data and Variable Selection for Achievement Estimation in Large‐Scale Assessments

  • Abstract
  • Literature Map
  • Similar Papers
Abstract
Translate article icon Translate Article Star icon

Abstract The operational scaling in large‐scale assessments (LSAs) comprises two stages: item response theory and latent regression modeling. Principal component analyses (PCA) are routinely performed before latent regression modeling for dimension reduction, but this approach retains too many principal components (PCs), threatening numerical stability. This study proposes a new approach: adding process variables to the usual contextual variables and replacing PCA with variable selection for latent regression modeling. We found that using Lasso, random forests, and ultimately gradient boosting for variable selection led to measurement precision similar to or higher than the PCA approach, but with considerably fewer covariates; including process variables into latent regression models consistently improved measurement precision. Integrating process data and variable selection yielded the highest measurement precision while achieving parsimony: The latent regression models with the 150 most important variables, including process variables, outperformed those with 330 PCs in France and 270 PCs in the Republic of Korea. The findings suggest that our proposed approach can effectively solve the overparameterization problem in LSA scaling while preserving measurement precision.

Similar Papers
  • PDF Download Icon
  • Research Article
  • Cite Count Icon 18
  • 10.3102/1076998620945199
Estimation of Latent Regression Item Response Theory Models Using a Second-Order Laplace Approximation
  • Aug 13, 2020
  • Journal of Educational and Behavioral Statistics
  • Björn Andersson + 1 more

The estimation of high-dimensional latent regression item response theory (IRT) models is difficult because of the need to approximate integrals in the likelihood function. Proposed solutions in the literature include using stochastic approximations, adaptive quadrature, and Laplace approximations. We propose using a second-order Laplace approximation of the likelihood to estimate IRT latent regression models with categorical observed variables and fixed covariates where all parameters are estimated simultaneously. The method applies when the IRT model has a simple structure, meaning that each observed variable loads on only one latent variable. Through simulations using a latent regression model with binary and ordinal observed variables, we show that the proposed method is a substantial improvement over the first-order Laplace approximation with respect to the bias. In addition, the approach is equally or more precise to alternative methods for estimation of multidimensional IRT models when the number of items per dimension is moderately high. Simultaneously, the method is highly computationally efficient in the high-dimensional settings investigated. The results imply that estimation of simple-structure IRT models with very high dimensions is feasible in practice and that the direct estimation of high-dimensional latent regression IRT models is tractable even with large sample sizes and large numbers of items.

  • Research Article
  • Cite Count Icon 43
  • 10.1111/jedm.12072
Extended Mixed‐Effects Item Response Models With the MH‐RM Algorithm
  • Jun 1, 2015
  • Journal of Educational Measurement
  • R Philip Chalmers

A mixed‐effects item response theory (IRT) model is presented as a logical extension of the generalized linear mixed‐effects modeling approach to formulating explanatory IRT models. Fixed and random coefficients in the extended model are estimated using a Metropolis‐Hastings Robbins‐Monro (MH‐RM) stochastic imputation algorithm to accommodate for increased dimensionality due to modeling multiple design‐ and trait‐based random effects. As a consequence of using this algorithm, more flexible explanatory IRT models, such as the multidimensional four‐parameter logistic model, are easily organized and efficiently estimated for unidimensional and multidimensional tests. Rasch versions of the linear latent trait and latent regression model, along with their extensions, are presented and discussed, Monte Carlo simulations are conducted to determine the efficiency of parameter recovery of the MH‐RM algorithm, and an empirical example using the extended mixed‐effects IRT model is presented.

  • Research Article
  • Cite Count Icon 46
  • 10.1177/00131644211045351
A Multilevel Mixture IRT Framework for Modeling Response Times as Predictors or Indicators of Response Engagement in IRT Models
  • Sep 13, 2021
  • Educational and Psychological Measurement
  • Gabriel Nagy + 1 more

Disengaged item responses pose a threat to the validity of the results provided by large-scale assessments. Several procedures for identifying disengaged responses on the basis of observed response times have been suggested, and item response theory (IRT) models for response engagement have been proposed. We outline that response time-based procedures for classifying response engagement and IRT models for response engagement are based on common ideas, and we propose the distinction between independent and dependent latent class IRT models. In all IRT models considered, response engagement is represented by an item-level latent class variable, but the models assume that response times either reflect or predict engagement. We summarize existing IRT models that belong to each group and extend them to increase their flexibility. Furthermore, we propose a flexible multilevel mixture IRT framework in which all IRT models can be estimated by means of marginal maximum likelihood. The framework is based on the widespread Mplus software, thereby making the procedure accessible to a broad audience. The procedures are illustrated on the basis of publicly available large-scale data. Our results show that the different IRT models for response engagement provided slightly different adjustments of item parameters of individuals’ proficiency estimates relative to a conventional IRT model.

  • Book Chapter
  • Cite Count Icon 2
  • 10.1007/978-3-030-05584-4_23
Applying the General Diagnostic Model to Proficiency Data from a National Skills Survey
  • Jan 1, 2019
  • Xueli Xu + 1 more

Large-scale educational surveys (including NAEP, TIMSS, PISA) utilize item-response-theory (IRT) calibration together with a latent regression model to make inferences about subgroup ability distributions, including subgroup means, percentiles, as well as standard deviations. It has long been recognized that grouping variables not included in the latent regression model can produce secondary bias in estimates of group differences (Mislevy, RJ, Psychometrika 56:177–196, 1991). To accommodate the ever-increasing number of background variables collected and required for reporting purposes, a principal component analysis based on the background variables (von Davier M, Sinharay S, Oranje A, Beaton AE, The statistical procedures used in national assessment of educational progress: recent developments and future directions. In: Rao CR, Sinharay S (eds) Handbook of statistics: vol. 26. Psychometrics. Elsevier B.V, Amsterdam, pp 1039–1055, 2007; Moran R, Dresher A, Results from NAEP marginal estimation research on multivariate scales. Paper presented at the annual meeting of the National Council on Measurement in Education, Chicago, 2007; Oranje A, Li D, On the role of background variables in large scale survey assessments. Paper presented at the annual meeting of the National Council on Measurement in Education, New York, NY, 2008) is utilized to keep the number of predictors in the latent regression models within a reasonable range. However, even this approach often results in the inclusion of several hundred variables, and it is unknown whether the principal component approach or similar approaches (such as latent-class approaches) are able to generate consistent estimates for individual subgroups (e.g., Wetzel E, Xu X, von Davier M, Educ Psychol Meas 75(5):1–25, 2014). The primary goal of the current study is to provide an exemplary application of diagnostic models for large-scale-assessment data. Specifically, a latent-class structure is used for covariates while continuing to use IRT models for item responses in the analytic model. Previous applications focused on adult literacy data (von Davier M, Yamamoto K, A class of models for cognitive diagnosis. Paper presented at the 4th Spearman invitational conference, Philadelphia, PA, 2004), as well as large-scale English-language testing programs (von Davier M; A general diagnostic model applied to language testing data (Research report no. RR-05-16). Educational Testing Service, Princeton, 2005, von Davier M, The mixture general diagnostic model. In: Hancock GR, Samuelsen KM (eds) Advances in latent variable mixture models. Information Age Publishing, Charlotte, pp 255–274, 2008), while the current application uses diagnostic modeling approaches on data from NAEP.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 46
  • 10.1186/s40536-016-0025-3
Assessment of fit of item response theory models used in large-scale educational survey assessments
  • Jul 8, 2016
  • Large-scale Assessments in Education
  • Peter W Van Rijn + 3 more

Latent regression models are used for score-reporting purposes in large-scale educational survey assessments such as the National Assessment of Educational Progress (NAEP) and Trends in International Mathematics and Science Study (TIMSS). One component of these models is based on item response theory. While there exists some research on assessment of fit of item response theory models in the context of large-scale assessments, there is a scope of further research on the topic. We suggest two types of residuals to assess the fit of item response theory models in the context of large-scale assessments. The Type I error rates and power of the residuals are computed from simulated data. The residuals are computed using data from four NAEP assessments. Misfit was found for all data sets for both types of residuals, but the practical significance of the misfit was minimal.

  • Research Article
  • Cite Count Icon 4
  • 10.1111/jedm.12348
Fully Gibbs Sampling Algorithms for Bayesian Variable Selection in Latent Regression Models
  • Oct 25, 2022
  • Journal of Educational Measurement
  • Kazuhiro Yamaguchi + 1 more

This study proposed Gibbs sampling algorithms for variable selection in a latent regression model under a unidimensional two‐parameter logistic item response theory model. Three types of shrinkage priors were employed to obtain shrinkage estimates: double‐exponential (i.e., Laplace), horseshoe, and horseshoe+ priors. These shrinkage priors were compared to a uniform prior case in both simulation and real data analysis. The simulation study revealed that two types of horseshoe priors had a smaller root mean square errors and shorter 95% credible interval lengths than double‐exponential or uniform priors. In addition, the horseshoe+ prior was slightly more stable than the horseshoe prior. The real data example successfully proved the utility of horseshoe and horseshoe+ priors in selecting effective predictive covariates for math achievement.

  • Research Article
  • Cite Count Icon 18
  • 10.1016/j.measurement.2018.01.044
Causal inference with latent variables from the Rasch model as outcomes
  • Feb 2, 2018
  • Measurement
  • Matthew P Rabbitt

Causal inference with latent variables from the Rasch model as outcomes

  • Research Article
  • Cite Count Icon 11
  • 10.1002/j.2333-8504.2004.tb01961.x
APPLICATION OF THE STOCHASTIC EM METHOD TO LATENT REGRESSION MODELS
  • Dec 1, 2004
  • ETS Research Report Series
  • Matthias Von Davier + 1 more

ABSTRACTThe reporting methods used in large scale assessments such as the National Assessment of Educational Progress (NAEP) rely on a latent regression model. The first component of the model consists of a p‐scale IRT measurement model that defines the response probabilities on a set of cognitive items in p scales depending on a p‐dimensional latent trait variable θ = (θ1, … θp). In the second component, the conditional distribution of this latent trait variable θ is modeled by a multivariate, multiple regression on a set of predictor variables, which are usually based on student, school and teacher variables in assessments such as NAEP.In order to fit the latent regression model using the maximum (marginal) likelihood estimation technique, multivariate integrals have to be evaluated. In the computer program MGROUP used by ETS for fitting the latent regression model to data from NAEP and other sources, the integration is currently done either by numerical quadrature (for problems up to two dimensions) or by an approximation of the integral. CGROUP, the current operational version of the MGROUP program used in NAEP and other assessments since 1993, is based on Laplace approximation that may not provide fully satisfactory results, especially if the number of items per scale is small.This paper examines the application of stochastic expectation‐maximization (EM) methods (where an integral is approximated by an average over a random sample) to NAEP‐like settings. We present a comparison of CGROUP with a promising implementation of the stochastic EM algorithm that utilizes importance sampling. Simulation studies and real data analysis show that the stochastic EM method provides a viable alternative to CGROUP for fitting multivariate latent regression models.

  • Research Article
  • Cite Count Icon 3
  • 10.1002/j.2333-8504.2009.tb02166.x
STOCHASTIC APPROXIMATION METHODS FOR LATENT REGRESSION ITEM RESPONSE MODELS
  • Jun 1, 2009
  • ETS Research Report Series
  • Matthias Von Davier + 1 more

ABSTRACTThis paper presents an application of a stochastic approximation EM‐algorithm using a Metropolis‐Hastings sampler to estimate the parameters of an item response latent regression model. Latent regression models are extensions of item response theory (IRT) to a 2‐level latent variable model in which covariates serve as predictors of the conditional distribution of ability. Applications for estimating latent regression models for data from the 2000 National Assessment of Educational Progress (NAEP) grade 4 math assessment and the 2002 grade 8 reading assessment are presented and results of the proposed method are compared to results obtained using current operational procedures.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 31
  • 10.3390/psych3030023
Estimating Explanatory Extensions of Dichotomous and Polytomous Rasch Models: The eirm Package in R
  • Jul 29, 2021
  • Psych
  • Okan Bulut + 2 more

Explanatory item response modeling (EIRM) enables researchers and practitioners to incorporate item and person properties into item response theory (IRT) models. Unlike traditional IRT models, explanatory IRT models can explain common variability stemming from the shared variance among item clusters and person groups. In this tutorial, we present the R package eirm, which provides a simple and easy-to-use set of tools for preparing data, estimating explanatory IRT models based on the Rasch family, extracting model output, and visualizing model results. We describe how functions in the eirm package can be used for estimating traditional IRT models (e.g., Rasch model, Partial Credit Model, and Rating Scale Model), item-explanatory models (i.e., Linear Logistic Test Model), and person-explanatory models (i.e., latent regression models) for both dichotomous and polytomous responses. In addition to demonstrating the general functionality of the eirm package, we also provide real-data examples with annotated R codes based on the Rosenberg Self-Esteem Scale.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 1
  • 10.5937/istrped2402377e
Analyzing immigrant and non-immigrant belonging experiences through IRT in large-scale assessments: Insights from Costa Rica
  • Jan 1, 2024
  • Research in Pedagogy
  • Elizondo-Gonzalez Fabian

Studies across countries part of the Organization for Economic Cooperation and Development (OECD) suggest that learners' sense of school belonging is often influenced by their place of birth. However, large-scale assessment studies rarely explore whether differences in belonging scores between nonimmigrant and immigrant learners are due to test bias. This study fills that gap by examining belonging scores in Costa Rican high schools using the Programme for International Student Assessment (PISA) 2022 data, with a focus on test fairness. Utilizing Multiple-group Differential Item Functioning (DIF) analyses and Item Response Theory (IRT) modeling, results show that all items are DIF-free, confirming no test bias. An Independent Samples t-test reveals no significant differences in belonging scores between immigrant and non-immigrant learners, which is a positive finding. It suggests that Costa Rican educational environments foster a shared sense of belonging, regardless of learners' place of birth. The most discriminating items, identified through IRT modeling, relate to performative or participatory aspects of school belonging. This study highlights the importance of incorporating IRT modeling and fairness protocols in large-scale assessments. By confirming that sense of belonging is not impacted by place of birth, stakeholders can confidently make decisions that further support inclusive educational practices in Costa Rican classrooms, knowing the PISA sense of belonging index provides unbiased, reliable scores.

  • Front Matter
  • Cite Count Icon 34
  • 10.1016/s1551-7144(09)00212-2
Classical and modern measurement theories, patient reports, and clinical outcomes
  • Jan 1, 2010
  • Contemporary Clinical Trials
  • Rochelle E Tractenberg

Classical and modern measurement theories, patient reports, and clinical outcomes

  • Research Article
  • Cite Count Icon 25
  • 10.3102/1076998607300422
An Importance Sampling EM Algorithm for Latent Regression Models
  • Sep 1, 2007
  • Journal of Educational and Behavioral Statistics
  • Matthias Von Davier + 1 more

Reporting methods used in large-scale assessments such as the National Assessment of Educational Progress (NAEP) rely on latent regression models. To fit the latent regression model using the maximum likelihood estimation technique, multivariate integrals must be evaluated. In the computer program MGROUP used by the Educational Testing Service for fitting the latent regression model to data from NAEP and other assessments, the integral is computed either by numerical quadrature or approximated. CGROUP, the current operational version of MGROUP used in NAEP for problems with more than two dimensions, uses Laplace approximation that may not provide fully satisfactory results, especially if the number of items per scale is small. This article examines a stochastic expectation-maximization (EM) method that uses importance sampling to NAEP-like settings. A simulation study and a real data analysis show that the importance sampling EM method provides a viable alternative to CGROUP for fitting multivariate latent regression models.

  • PDF Download Icon
  • Research Article
  • Cite Count Icon 28
  • 10.3389/fpsyg.2017.00484
Practical Consequences of Item Response Theory Model Misfit in the Context of Test Equating with Mixed-Format Test Data
  • Apr 4, 2017
  • Frontiers in Psychology
  • Yue Zhao + 1 more

In item response theory (IRT) models, assessing model-data fit is an essential step in IRT calibration. While no general agreement has ever been reached on the best methods or approaches to use for detecting misfit, perhaps the more important comment based upon the research findings is that rarely does the research evaluate IRT misfit by focusing on the practical consequences of misfit. The study investigated the practical consequences of IRT model misfit in examining the equating performance and the classification of examinees into performance categories in a simulation study that mimics a typical large-scale statewide assessment program with mixed-format test data. The simulation study was implemented by varying three factors, including choice of IRT model, amount of growth/change of examinees’ abilities between two adjacent administration years, and choice of IRT scaling methods. Findings indicated that the extent of significant consequences of model misfit varied over the choice of model and IRT scaling methods. In comparison with mean/sigma (MS) and Stocking and Lord characteristic curve (SL) methods, separate calibration with linking and fixed common item parameter (FCIP) procedure was more sensitive to model misfit and more robust against various amounts of ability shifts between two adjacent administrations regardless of model fit. SL was generally the least sensitive to model misfit in recovering equating conversion and MS was the least robust against ability shifts in recovering the equating conversion when a substantial degree of misfit was present. The key messages from the study are that practical ways are available to study model fit, and, model fit or misfit can have consequences that should be considered when choosing an IRT model. Not only does the study address the consequences of IRT model misfit, but also it is our hope to help researchers and practitioners find practical ways to study model fit and to investigate the validity of particular IRT models for achieving a specified purpose, to assure that the successful use of the IRT models are realized, and to improve the applications of IRT models with educational and psychological test data.

  • Research Article
  • Cite Count Icon 8
  • 10.1177/00131644221111838
Assessing Dimensionality of IRT Models Using Traditional and Revised Parallel Analyses.
  • Jul 21, 2022
  • Educational and psychological measurement
  • Wenjing Guo + 1 more

Determining the number of dimensions is extremely important in applying item response theory (IRT) models to data. Traditional and revised parallel analyses have been proposed within the factor analysis framework, and both have shown some promise in assessing dimensionality. However, their performance in the IRT framework has not been systematically investigated. Therefore, we evaluated the accuracy of traditional and revised parallel analyses for determining the number of underlying dimensions in the IRT framework by conducting simulation studies. Six data generation factors were manipulated: number of observations, test length, type of generation models, number of dimensions, correlations between dimensions, and item discrimination. Results indicated that (a) when the generated IRT model is unidimensional, across all simulation conditions, traditional parallel analysis using principal component analysis and tetrachoric correlation performs best; (b) when the generated IRT model is multidimensional, traditional parallel analysis using principal component analysis and tetrachoric correlation yields the highest proportion of accurately identified underlying dimensions across all factors, except when the correlation between dimensions is 0.8 or the item discrimination is low; and (c) under a few combinations of simulated factors, none of the eight methods performed well (e.g., when the generation model is three-dimensional 3PL, the item discrimination is low, and the correlation between dimensions is 0.8).

Save Icon
Up Arrow
Open/Close
Notes

Save Important notes in documents

Highlight text to save as a note, or write notes directly

You can also access these Documents in Paperpal, our AI writing tool

Powered by our AI Writing Assistant