Assessing the Unconditional and Conditional External Validity of Noncognitive Test Scores: A Unifying Model-Based Proposal.
Evidence of external validity based on individual score estimates is still relevant in many psychometric applications. From a model-based perspective, however, the topic appears to have been rather neglected in recent decades. Thus, in structural equation modelling (SEM), this evidence is sought to be obtained structurally, bypassing the scoring stage. And, in item response theory (IRT), the score interest mostly focuses on internal properties. Taking this state of affairs into account, this paper develops and proposes a model-based approach, intended for noncognitive measures, that combines SEM and IRT developments, and which allows a detailed assessment of the external validity of a class of score estimates to be carried out. The starting point is a general extended model that also includes the relevant external variables. From this general model, four well-known extended IRT models can be derived and fitted at the structural level. Next, on the basis of the structural results, a series of unconditional (population-dependent) and conditional (population-independent) indices that describe the model-implied relation between the score estimates and each external variable are developed and proposed. The practical relevance of the proposal is discussed mainly around three applications: assessing model appropriateness, obtaining point and interval prediction estimates at the individual level, and shortening a test while optimizing the external validity of the resulting version. The functioning of the proposal is illustrated using a real-data example.
- Front Matter
34
- 10.1016/s1551-7144(09)00212-2
- Jan 1, 2010
- Contemporary Clinical Trials
Classical and modern measurement theories, patient reports, and clinical outcomes
- Research Article
8
- 10.1177/0962280213504177
- Jul 11, 2016
- Statistical Methods in Medical Research
Both item response theory and structural equation models are useful in the analysis of ordered categorical responses from health assessment questionnaires. We highlight the advantages and disadvantages of the item response theory and structural equation modelling approaches to modelling ordinal data, from within a community health setting. Using data from the SPARCLE project focussing on children with cerebral palsy, this paper investigates the relationship between two ordinal rating scales, the KIDSCREEN, which measures quality-of-life, and Life-H, which measures participation. Practical issues relating to fitting models, such as non-positive definite observed or fitted correlation matrices, and approaches to assessing model fit are discussed. item response theory models allow properties such as the conditional independence of particular domains of a measurement instrument to be assessed. When, as with the SPARCLE data, the latent traits are multidimensional, structural equation models generally provide a much more convenient modelling framework.
- Research Article
- 10.59863/optz4045
- Dec 1, 2022
- Chinese/English Journal of Educational Measurement and Evaluation
This essay sketches the historical development of latent variable scoring procedures in the item response theory (IRT) and factor analysis literatures, observing that the most commonly used score estimates in both traditions are fundamentally the same; only methods of calculation differ. Different procedures have been used to derive factor score estimates and latent variable estimates in IRT, and different computational procedures have been the result. Due to differences in the context of score usage, challenges have led to different solutions in the IRT and factor analytic traditions. The needs for bias corrections differ, as do the corrections that have been proposed. While the standard factor analysis model has naturally Gaussian likelihoods, IRT does not, but in IRT normal approximations have been used in various contexts to make the IRT computations more like those of factor analysis. Finally, factor analysis alone has been the home of decades of controversy over factor score indeterminacy, while IRT has not, even though the scores in question are the same. That is an artifact of history and the ways the models have been written in the IRT and factor analytic literatures. IRT has never been plagued with questions of indeterminacy, which helps to clarify the position that what is referred to as indeterminacy is not a problem.
- Dissertation
- 10.17077/etd.9yfctk57
- Feb 5, 2015
<p>The use of testlets in a test can cause multidimensionality and local item dependence (LID), which can result in inaccurate estimation of item parameters, and in turn compromise the quality of item response theory (IRT) true and observed score equating of testlet-based tests. Both unidimensional and multidimensional IRT models have been developed to control local item dependence caused by testlets. The purposes of the current study were to (1) investigate how different levels of LID can affect IRT true and observed score equating of testlet-based tests when the traditional three parameter logistic (3PL) IRT model was used for calibration, and (2) compare the performance of four different IRT models, including the 3PL IRT model, graded response model (GRM), testlet response theory model (TRT), and bifactor model, in IRT true and observed score equating of testlet-based tests with various levels of local item dependence.</p><p>Both real and simulated data analyses were conducted in this study. Two testlet-based tests (i.e., Test A and Test B) that differed in subjects, test length, and testlet length were used in the real data analysis. For simulated data analysis, two main factors were investigated in this study: (1) testlet length (5 or 10), and (2) LID level within testlets that was defined by testlet effect variance (0, 0.25, 0.5625, 0.75, 1, and 1.5). For the unidimensional IRT models (i.e., 3PL IRT model and GRM), unidimensional IRT true score and observed score equating procedures, explained in Kolen and Brennan (2004), were used. For the two investigated multidimensional IRT models (i.e., 3PL TRT model and bifactor model), the unidimensional approximation of multidimensional item response theory (MIRT) true score equating procedure and the unidimensional approximation of MIRT observed score equating procedure (Brossman & Lee, 2013) were applied. The traditional equipercentile equating method was used as the baseline for comparison in both real data and simulated data analyses.</p><p>It was found in the study that both testlet length and the LID level affected the performance of the investigated models on IRT true and observed score equating of testlet-based tests. When the traditional 3PL IRT model was used for tests with long testlets, higher levels of local item dependence led to IRT equating results that deviated further away from those obtained from the baseline method. However, the effect of local item dependence on IRT equating results was not prominent for tests with short testlets.</p><p>Moreover, for tests consisting of long testlets (e.g., a testlet length of 10 or more) and having a very low level of local item dependence (e.g., a LID level of 0.25 or lower), and for tests consisting of short testlets (e.g., a testlet length around 5), all four investigated IRT models worked well in IRT true and observed score equating. For tests with long testlets and a relatively high level of local item dependence (e.g., a LID level of 0.5625 or higher), the GRM, bifactor, and TRT models outperformed the traditional 3PL IRT model in IRT true and observed equating of testlet-based tests.</p><p>The study suggested that the selection of models for IRT true and observed score equating of testlet-based tests should be considered with respect to the features of the testlet-based tests and the groups of examinees from which the data is collected. It is hoped that this study encourages researchers to identify differences among existing models for IRT true and observed score equating of testlet-based tests with various features, and to develop new models that are appropriate for modeling testlet-based tests to obtain accurate IRT number correct score equating results.</p>
- Book Chapter
14
- 10.1007/978-981-10-3302-5_5
- Jan 1, 2016
Classical Test Theory (CTT), also known as the true score theory, refers to the analysis of test results based on test scores. The statistics produced under CTT include measures of item difficulty, item discrimination, measurement error and test reliability. The term “Classical” is used in contrast to “Modern” test theory which usually refers to item response theory (IRT). The fact that CTT was developed before IRT does not mean that CTT is outdated or replaced by IRT. Both CTT and IRT provide useful statistics to help us analyse test data. Generally, CTT and IRT provide complementary results. For many item analyses, CTT may be sufficient to provide the information we need. There are, however, theoretical differences between CTT and IRT, and many researchers prefer IRT because of enhanced measurement properties under IRT. IRT also provides a framework that facilitates test equating, computer adaptive testing and test score interpretation. While this book devotes a large part to IRT, we stress that CTT is an important part of the methodologies for educational and psychological measurement. In particular, the exposition of the concept of reliability in CTT sets the basis for evaluating measuring instruments. A good understanding of CTT lays the foundations for measurement principles. There are other approaches to measurement such as generalizability theory and structural equation modelling, but these are not the focus of attention in this book.
- Conference Article
- 10.1063/1.4992683
- Jan 1, 2017
- AIP conference proceedings
The Item Response Theory (IRT) has become one of the most popular scoring frameworks for measurement data, frequently used in computerized adaptive testing, cognitively diagnostic assessment and test equating. According to Andrade et al. (2000), IRT can be defined as a set of mathematical models (Item Response Models – IRM) constructed to represent the probability of an individual giving the right answer to an item of a particular test. The number of Item Responsible Models available to measurement analysis has increased considerably in the last fifteen years due to increasing computer power and due to a demand for accuracy and more meaningful inferences grounded in complex data. The developments in modeling with Item Response Theory were related with developments in estimation theory, most remarkably Bayesian estimation with Markov chain Monte Carlo algorithms (Patz & Junker, 1999). The popularity of Item Response Theory has also implied numerous overviews in books and journals, and many connections between IRT and other statistical estimation procedures, such as factor analysis and structural equation modeling, have been made repeatedly (Van der Lindem & Hambleton, 1997). As stated before the Item Response Theory covers a variety of measurement models, ranging from basic one-dimensional models for dichotomously and polytomously scored items and their multidimensional analogues to models that incorporate information about cognitive sub-processes which influence the overall item response process. The aim of this work is to introduce the main concepts associated with one-dimensional models of Item Response Theory, to specify the logistic models with one, two and three parameters, to discuss some properties of these models and to present the main estimation procedures.
- Research Article
- 10.1177/0272989x251340990
- Jun 13, 2025
- Medical decision making : an international journal of the Society for Medical Decision Making
ObjectivesThe EQ-5D-5L and Patient-Reported Outcomes Measurement Information System (PROMIS®) preference score (PROPr) are preference-based measures. This study compares mapping and linking approaches to align the PROPr and the PROMIS domains included in PROPr plus Anxiety with EQ-5D-5L item responses and preference scores.MethodsA general population sample of 983 adults completed the online survey. Regression-based mapping methods and item response theory (IRT) linking methods were used to align scores. Mapping was used to predict EQ-5D-5L item responses or preference scores using PROMIS domain scores. Equating strategies were applied to address regression to the mean. The linking approach estimated item parameters of EQ-5D-5L based on the PROMIS score metric and generated bidirectional crosswalks between EQ-5D-5L item responses and relevant PROMIS domain scores.ResultsEQ-5D-5L item responses were significantly accounted for by PROMIS domains of Anxiety, Depression, Fatigue, Pain Interference, Physical Function, Social Roles, and Sleep Disturbance. EQ-5D-5L preference scores were accounted for by the same PROMIS domains, excluding Anxiety and Fatigue, and by the PROPr preference scores. IRT-linking crosswalks were generated between EQ-5D-5L item responses and PROMIS domains of Physical Function, Pain, and Depression. Small differences were found between observed and predicted scores for all 3 methods. The direct mapping approach (directly predicting EQ-5D-5L scores) with the equipercentile equating strategy proved superior to the linking method due to improved prediction accuracy and comparable score range coverage.ConclusionsThe PROPr and the PROMIS domains included in the PROMIS-29+2 predict EQ-5D-5L preference scores or item responses. Both methods can generate acceptably precise EQ-5D-5L preference scores, with the direct mapping approach using the equating strategy offering better precision. We summarized recommended score conversion tables based on available and desired scores.HighlightsThis study compares mapping (score prediction) and IRT-based linking approaches to align the PROPr and the PROMIS domains with EQ-5D-5L item responses and preference scores.Researchers, clinicians, and stakeholders can use this study's regression formulas and score crosswalks to convert scores between PROMIS and EQ-5D-5L.Mapping can generate more precise scores, while linking offers greater flexibility in score estimation when fewer PROMIS domain scores are collected.
- Research Article
1
- 10.24230/kjiop.v25i2.421-452
- May 31, 2012
- Korean Journal of Industrial and Organizational Psychology
The present study investigated the utilities of two types of item response process models(dominance model and ideal point model) for personality item parameter estimation and scoring. The authors developed scales for four personality traits(achievement, fairness, cooperation and honesty) using classical test theory, dominance item response theory(IRT) method, and ideal point IRT method and compared the methods in terms of model-data fit, information and criterion validity. Results show that the fit of ideal point IRT model was better than that of dominance IRT model, but the difference between the fit of two models was very slight. The test information functions of ideal point IRT model and dominance IRT model for honesty and cooperation scales were very similar. The criterion-related validity based on individual ability estimates and grades was not significant for the three methods but the validity for the ideal point method is not better than dominant IRT model. Implications and limitations of the findings are discussed.
- Research Article
20
- 10.3102/1076998607306451
- Dec 1, 2008
- Journal of Educational and Behavioral Statistics
The randomized response technique ensures that individual item responses, denoted as true item responses, are randomized before observing them and so-called randomized item responses are observed. A relationship is specified between randomized item response data and true item response data. True item response data are modeled with a (non)linear mixed effects and/or item response theory model. Although the individual true item responses are masked through randomizing the responses, the model extension enables the computation of individual true item response probabilities and estimates of individuals’ sensitive behavior/attitude and their relationships with background variables taking into account any clustering of respondents. Results are presented from a College Alcohol Problem Scale (CAPS) where students were interviewed via direct questioning or via a randomized response technique. A Markov Chain Monte Carlo algorithm is given for estimating simultaneously all model parameters given hierarchical structured binary or polytomous randomized item response data and background variables.
- Research Article
25
- 10.1027/1015-5759/a000609
- Jul 1, 2020
- European Journal of Psychological Assessment
When constructing a questionnaire to assess a psychological construct, one important decision researchers have to make is how to collect responses from test takers; that is, which response format to implement.We argued in a previous editorial published in the European Journal of Psychological Assessment (EJPA) that this decision deserves more attention and should be an explicit step in the test construction process (Wetzel & Greiff, 2018).The reason for this is that it can be a consequential decision that influences the validity of conclusions we draw about test takers' trait levels or about relations between constructs and criteria (Brown & Maydeu-Olivares, 2013; Wetzel & Frick, 2020).In this editorial, which can be considered a followup to the first one, we will take a closer look at two response formats 1 : rating scales (RS), the current default in most questionnaires, and the multidimensional forced-choice (MFC) format, an alternative that is currently the focus of a considerable body of research.We will first define the two formats and point out some of their advantages and disadvantages.Then, we will provide a summary and evaluation of research comparing RS and MFC.Third, we will draw some preliminary conclusions on the feasibility of applying MFC as an alternative to RS. Fourth, we will point out some open research questions.We will end with some recommendations and implications for readers and authors of EJPA.In this editorial, the overall goal is to give researchers and test users an overview of the current state of the research on RS versus MFC and to provide guidance on the feasibility of applying MFC in research on psychological assessment.1 The multidimensional forced-choice format is both an item and a response format.For simplicity in the comparison with rating scales, we refer to it as response format.
- Research Article
- 10.3724/sp.j.1041.2008.00092
- Nov 25, 2008
- Acta Psychologica Sinica
Both item response theory(IRT) and a multilevel model are used in a variety of social science research applications.The use of IRT allows researchers to link the observed categorical responses provided by students with an underlying unobservable trait,such as ability or attitude.A multilevel model allows the natural multilevel structure that is widely present in social science data to be represented formally in data analysis.In some cases,researchers may wish to study the effects of the covariates on the latent trait.These covariates may include information pertaining to responses as well as contextual information.Traditionally,manifest variables are used in a multilevel analysis as fixed and known entities.An important deficiency is that the measurement error associated with the test scores is ignored.In general,the use of unreliable test scores leads to a biased estimation of the regression coefficients;consequently,the resulting statistical inference can be rather misleading. In this paper,a multilevel IRT model is presented where some of the variables cannot be observed directly but are measured using tests or questionnaires.A two-level model can be defined using a two-level formulation in which level 1 is the item-level model and level 2 is the person-level model.In such a model,the latent ability and item parameters can be estimated simultaneously.Further,we can consider the effects of the person-level covariates on the person-level ability.Similar to a traditional two-level model,the two-level IRT model can be expanded to a three-level model,which includes a higher site level.The threelevel IRT model proposed in this paper yields the additional benefit of being able to accommodate data that are collected in a hierarchical setting.This expansion of multilevel IRT models to three levels allows not only the dependency typically found in hierarchical data to be accommodated but also the estimation of latent traits at different levels as well as the estimation of relationships between predictor variables and latent traits at different levels. The purpose of this paper is to provide both a theoretic description and a practical application of the multilevel IRT model.First,we demonstrate the algebraic equivalence between the parameterizations of an IRT model and the multilevel IRT model.Second,we illustrate the method,using an application that involves students' achievements on a mathematics test and test results regarding the characteristics of students and schools.Third,we expand the multilevel IRT model to the identification of differential item functioning(DIF).Finally,we present a discussion on the advantages and disadvantages of using multilevel IRT models in applied research and provide the scope for future research.
- Research Article
13
- 10.1177/0013164417745949
- Dec 6, 2017
- Educational and Psychological Measurement
This note highlights and illustrates the links between item response theory and classical test theory in the context of polytomous items. An item response modeling procedure is discussed that can be used for point and interval estimation of the individual true score on any item in a measuring instrument or item set following the popular and widely applicable graded response model. The method contributes to the body of research on the relationships between classical test theory and item response theory and is illustrated on empirical data.
- Research Article
13
- 10.1093/swr/34.2.94
- Jun 1, 2010
- Social Work Research
The need to develop measures that tap into constructs of interest to social work, refine existing measures, and ensure that measures function adequately across diverse populations of interest is critical. Item response theory (IRT) is a modern measurement approach that is increasingly seen as an essential tool in a number of allied professions. IRT-based measurement uses a model-based approach that has several analytical and explanatory advantages over classical test theory. In particular, IRT-based techniques facilitate the process of specific item selection, allow for increased measurement precision with fewer items, and provide greater capacity for understanding and accounting for measurement bias across diverse populations. A survey of the top (as rated by impact factor) 20 social work journals revealed that few measurement articles in the social work literature use IRT or other modern measurement approaches. The benefit of incorporating more IRT-based approaches for developing, refining, and ensuring the application of measures to diverse populations is discussed. KEY WORDS: bias; classical test theory; item response theory; measurement; social work ********** The state of measurement within the social work literature is integrally related to knowledge base development and, ultimately, the extent to which research is able to meaningfully inform practice (Holden, Nizza, & Weissman, 1995). Scholarship highlights at least three measurement-related research domains within the field of social work. The first concerns the development of valid and reliable measures that capture the diverse set of phenomena relevant to social work, particularly those phenomena that may not be adequately represented by existent standardized instruments. The second is the assessment and validation of such measures. In particular, high-quality intervention research hinges on the validity and reliability of measures used to assess outcomes (Rosen, Proctor, & Staudt, 1999).Third, a growing body of literature challenges the extent to which well-validated measures adequately account and adjust for within- and across-population sources of diversity (see Ramirez, Ford, Stewart, & Teresi, 2005; Snowden, 2003), and such concerns are highly salient to social work's commitment to diversity-sensitive and -responsive research and practice. During the 1980s and 1990s, social work researchers outlined the relative benefits of item response theory (IRT) over classical test theory (CTT) measurement models, calling explicitly for IRT-based models' increased utilization to address measurement problems in social work research (DeRoos & Allen-Meares, 1993, 1998; Nugent & Hankins, 1989,1992). Indeed, IRT models have largely subsumed CTT approaches within a wide range of allied fields and disciplines (for example, medicine, psychology, nursing, public health, education) (see Dunn, Resnicow, & Klesges, 2006; Embretson & Reise, 2000; Fries, Bruce, & Cella, 2005; Lord, 1980; Ware, Bjorner, & Kosinski, 2000). Given early interest among social work researchers and the recent proliferation of IRT methods within other applied social sciences, our overall objective in the present study was to assess the extent to which these methods are represented within social work research. This review thus realizes three overlapping aims. First, it provides a description and comparison of IRT and CTT models and outlines the potential contributions of IRT methods to social work scholarship; it also briefly discusses IRT more generally as a latent variable model and its overlap with confirmatory factor analytic (CFA) and multi-level modeling methods. Second, it presents the results of a structured review assessing the penetration of IRT-based methods into the field of social work as reflected in key social work research journals. Third, using these results as a launching point, we highlight particular lines of inquiry within social work research where the application of IRT methods would likely yield substantial innovation. …
- Research Article
22
- 10.1177/0962280219884574
- Oct 30, 2019
- Statistical Methods in Medical Research
When assessing change in patient-reported outcomes, the meaning in patients’ self-evaluations of the target construct is likely to change over time. Therefore, methods evaluating longitudinal measurement non-invariance or response shift at item-level were proposed, based on structural equation modelling or on item response theory. Methods coming from Rasch measurement theory could also be valuable. The lack of evaluation of these approaches prevents determining the best strategy to adopt. A simulation study was performed to compare and evaluate the performance of structural equation modelling, item response theory and Rasch measurement theory approaches for item-level response shift detection. Performances of these three methods in different situations were evaluated with the rate of false detection of response shift (when response shift was not simulated) and the rate of correct response shift detection (when response shift was simulated). The Rasch measurement theory-based method performs better than the structural equation modelling and item response theory-based methods when recalibration was simulated. Consequently, the Rasch measurement theory-based approach should be preferred for studies investigating only recalibration response shift at item-level. For structural equation modelling and item response theory, the low rates of reprioritization detection raise issues on the potential different meaning and interpretation of reprioritization at item-level.
- Research Article
- 10.1007/s11205-011-9928-0
- Sep 10, 2011
- Social Indicators Research
A crucial issue in the European Union (EU) is which policies should be regulated by EU and which ones by national governments. Given this situation it is interesting to study the citizens’ preference for the level of political decision making. The interest of the paper is mainly empirical, which consists in the creation of a measure for supranationalism decision making from different measured policies and to study how personal variables affect the level of political decision. A combination of Item response theory (IRT) and structural equation modeling (SEM) will be used. IRT is useful to discover whether a cumulative scale for different policies exists, and SEM will be used as validity of the supranationalism scale. Multiple group Confirmatory Factor analysis taking into account invariance tests will be used to compare the differences within gender, age, education and political trust with supranationalism scale. A whole model combining all these predictors and the supranationalism scale can be done using multiple input multiple cause-MIMIC models. The application is based on Spanish data from the European Social Survey. Differences in education and age on the level of political decision making are found. The combination of IRT and SEM methods it of great interest in the case is the researcher wants to fulfill both goals, creating the construct for supranationalism and validation of the model.